On authenticity, imperfection, and what we’re actually looking for
Something has been added to the act of reading, and the addition has cost something. It’s happened so gradually that most people haven’t named it yet. But it’s there: a prior question, installed before the reading begins, that used to not be there. Not what is this saying — that question still comes, but later. First, now: who made this.
It arrives before you decide to ask it. Before anything specific flags it. Before the sentence has had a chance to do anything to you. The question is already running.
The project of identifying AI-generated text has developed its own forensic vocabulary. The em dash that arrives with slightly too much confidence. The “not X but Y” construction doing its small rhetorical pivot. Certain verbs — delve, resonate, unpack — clustering around a specific register of performed seriousness. Sentences that locate their point without much visible uncertainty about whether the point was worth locating.
I wrote that paragraph and read it back. It contains an em dash. The last sentence arrives at its point without much visible uncertainty about whether the point was worth the trip. I left it. Not as a demonstration — because I don’t know how to write about this without writing in it. Whatever habits I’ve absorbed from the critical tradition I came up reading, they’re inside me now, not outside me where I could set them down.
These features didn’t come from language models. They predate them by decades, some by centuries. The em dash belongs to the literary essayist. The “not X but Y” construction runs through the critical tradition of the last hundred years. Words like nuanced and tapestry — the whole vocabulary of polished approximation — were the professional language of editors and culture writers who reached for them without irony, because they signaled a certain kind of seriousness. Any writer who learned to write by reading serious criticism absorbed these habits, because that was how serious criticism moved.
Which means the forensic project of AI detection has become, without anyone intending it, a forensic project aimed at writers who learned to write from exactly the sources now treated as evidence. The em dash. The balanced construction. The careful vocabulary. All of it suddenly reading as suspect — not because it changed, but because the population of things that produce it did.
Here is what I think the detection project is actually trying to find, underneath the vocabulary of tells.
It isn’t AI. It’s the absence of evidence that someone was changed by writing this.
That’s the real tell, if there is one: not the em dash but the sense that the prose came out of a process that didn’t cost anyone anything. That the argument arrived pre-formed rather than worked toward. That the stumbles you’d expect from a mind genuinely working something out aren’t there — not because they were edited out, but because they never occurred.
This is a more interesting version of the detection question. It’s also harder, because it’s not actually a detection question at all. You can detect an em dash. You can’t detect the presence or absence of genuine cognitive effort from the surface of a text. Published writing was always already a cleaned-up version of the thinking behind it. The rough draft had the stumbles; the published essay had them removed. We’ve been reading reconstructions of thought and calling them thought, for as long as publishing has existed.
Early language models produced writing that was, in a specific way, too good. Too consistent. Arguments that resolved without remainder, sentences well-formed without exception — prose with the quality of a room prepared for visitors, nothing left out of place, no sign anyone had actually lived there. That absence of mess was itself the tell. Human writing, even careful human writing, isn’t like that. It has remainder. The most memorable sentences are often the ones that arrive somewhere the writer didn’t predict.
The models are learning this — not from a list of imperfections to imitate, but from enough exposure to human language that something resembling the texture of human unpredictability starts to emerge on its own. The better outputs now have a studied ambiguity to them. A deliberate residue of the inexact. A clause that doesn’t quite close when you’d expect. Under ordinary attention, this reads like the trace of a mind in motion.
What’s unsettling isn’t that this is technically impressive. It’s what it implies. The models aren’t learning to write like good writers write. They’re learning to produce the texture of a mind that hasn’t finished — rendered precisely enough that “hasn’t finished” starts to look like a style, available on request, rather than a state that has to be lived through.
We’ve never had clean access to the difference between actually thinking something through and producing the impression that thinking-through occurred. So this isn’t a technology problem. It’s a problem about what we mean when we say a sentence feels alive.
I spent years in an industry where the question of quality was settled by the room. Not by credentials, not by origin, not by who made it. By the room.
You read the temperature of a space, you made a decision, and within thirty seconds you knew whether you were right. The response was immediate and unambiguous. Nobody ran a verification check before deciding whether to move to something. The question was always: does this work? The answer was always already in the bodies in front of you.
What the room was asking — I’ve been thinking about this — is the same thing the detection tool is trying to ask. But the room asked it directly, and the tool is asking it by proxy. The room wanted to know: is something real moving through this? Does it carry something that arrives in you, changes the air, moves you from one state to another? The tool looks for evidence of that reality in the text’s provenance rather than its effect.
I’ve been in both of these worlds, and the difference is not small. In one, the question is answered by your own response. In the other, the question is whether your response can be trusted at all — whether the thing that moved you, or didn’t, might have been manufactured, and whether that changes what the movement was worth.
That replacement — of direct response with forensic suspicion — is what I keep wanting to name, and finding hard to name without it sounding smaller than it is.
In Go, after AlphaGo, there was a score. There was a winner. The field reorganized around what the machine had discovered, because the game had always had an external standard — the board, the territory, the count. The players didn’t want this. They hadn’t imagined it. But the game’s structure meant they couldn’t argue against it. The score said what it said.
Writing doesn’t have a score. There is no board, no count, no external measure that settles it. The room I’m describing from those years had something like a score — the response was immediate and measurable, in its way. Writing doesn’t even have that. What writing has is the judgment of whoever is reading it, and that judgment is slow, individual, impossible to aggregate.
Which means the detection obsession is attempting something the field doesn’t structurally support: a binary where none exists. If we can’t say this is better, we can say this is human. If we can’t say this failed, we can say this was generated.
The detection question is what you reach for when you can’t make the judgment the situation is actually asking for. Or when you’ve stopped believing the judgment is yours to make.
There are legitimate versions of this concern. In education they’re pressing. The argument for students working through difficult texts without automatic summarization — for essays that require a student to not-know their way toward understanding — that argument is sound. The process is the point. The stumble is the lesson. A mind that never has to stay uncertain for long is a different kind of mind from one that does.
But this concern doesn’t transfer cleanly to every reading situation. The professor reading a class submission and the reader encountering an essay in a magazine have different relationships to the process behind the text. The class exists to develop capacity. The magazine exists to deliver thought. The detection obsession has imported the educational concern into every reading situation, as though every text were a credential and every essay a proof of labor, without asking whether the import makes sense.
The question the detection obsession has displaced is the one that was there before it: what is good writing? What does it do? How does it move? Does it leave you somewhere different from where it found you?
I’m not sure why we’ve stopped asking this. The answer is harder to agree on, harder to operationalize, doesn’t produce a number. But here is what I think the detection question actually reveals about itself: it is what you ask when you’ve decided your response to a text can’t be trusted. When you believe the text might fool you. When your own attention, your history with language, your capacity to be moved or not moved — none of it is reliable enough to make the call.
That distrust is understandable. The situation that produced it is real. But I’m not sure that distrust is something to cultivate. In the room I’m describing, the only thing that mattered was whether you trusted your own ear. You could be wrong. You were wrong, sometimes. But the alternative — running a check before you let yourself respond — wasn’t available. The question came first and the answer came in the body, and then you acted.
The essay that arrives and makes something happen in you has already answered the question the detection tool is trying to ask. The answer is in what happened. Whether you let that answer stand, or whether you reach for the tool first — that’s a choice about what kind of reader you want to be.
I wrote several versions of that last sentence. Each time it arrived somewhere more conclusive than I intended, more like a position I’d prepared than one I’d found. This is the version that still has a question in it.
I’m not sure the question is comfortable. I’m leaving it there.