Why "you can always spot AI writing" tells you almost nothing about how much AI writing you have read.
"It is so easy to spot AI generated content. It always sounds the same."
I hear some version of this most weeks. It came up again recently, typed into a chat window during a presentation about AI, by somebody entirely confident about it.
And they are right. Anyone who reads a reasonable amount has built a working instinct for it. The particular cadence. The list of three. The paragraph that restates the previous paragraph in slightly different words. The closing sentence that gestures at significance without adding anything. The em dashes and the en dashes. Once you have seen it a few times you cannot unsee it.
So the observation is sound. It is the conclusion drawn from it that does not follow.
"You are not judging a sample of AI writing. You are judging a sample of AI writing that failed."
Every piece you have ever correctly identified has one property in common, and it is the property that put it in front of you.
I am probably worse at this than I think, in both directions. I have flagged writing as machine made that almost certainly was not, and I never find out about those either. Neither kind of error reports back.
The pale ones are not absent from the world. They are absent from the sample you formed the judgement on, and they are absent for exactly the reason you would want to know about.
There is a well known version of this error, and it is worth walking through, because everybody involved in it was clever and looking carefully at real data.
During the Second World War, analysts examined bombers returning from missions over Europe and mapped where they were taking damage. A clear pattern emerged: wings, fuselage, and the area around the tail gunner. The sensible response was to add armour where the bullet holes were.
Abraham Wald, a statistician working for the Statistical Research Group in New York, pointed out that they were examining the wrong aeroplanes. The map showed where a bomber could be hit and still get home. The aircraft hit in the engines were not in the sample. They had not come back, and neither had their crews.
The armour went on the engines.
Damage recorded on the aircraft that came back. The engines are not undamaged. They are unrepresented.
Now apply the same reading to your instinct about AI writing, or about any other kind of content if that is your domain.
Every piece you have correctly identified failed. It carried the tells, nobody removed them, and it reached you in that state. That is not incidental to your sample. It is the entry requirement.
The pieces that were carefully specified, reviewed, corrected and rewritten before publication did not announce themselves either. If the work was done properly you read them and thought nothing at all, which is the outcome the work was for. They are not in your sample and they cannot be, because the thing that would have placed them there is precisely the thing that was removed.
The error is not a failure of intelligence in either case. It is that absence does not announce itself. A sample that has been filtered feels exactly like a sample that has not, from the inside.
There is a version of this worth saying plainly, because it changes what the complaint is actually about.
When somebody says AI writing is obvious and bad, the observation underneath is accurate. What they encountered was obvious and bad. But the property being detected is not "written with AI." It is "published without review."
Those two things correlate strongly at the moment, which is why the instinct works as well as it does. They are not the same property, and the correlation is a fact about current habits rather than about the tool.
A first draft from a language model, shipped unread, has a signature. So does a first draft from a person, shipped unread. The difference is that nobody has built a folk theory about the second one, because it has been happening for as long as there has been writing, and we have other words for it.
I use AI extensively. Most days, across engineering work, client documents, code and writing. I am not a neutral party here and there is no point pretending otherwise.
What I have found is that the difference between output carrying the tells and output not carrying them is almost never the model. It is whether somebody specified what they actually wanted, and then checked whether they got it.
That still relies, for now, on a person being able to carry the specification. It is a partnership rather than a handover, and the responsibility sits on the human side of it.
Spell check has been on nearly every machine most of us have used for thirty years and it has never been perfect. Some of that is the tool. A good deal of it is somebody leaving the dictionary set to the wrong variant of English and then blaming the red underline.
None of the rules alongside are sophisticated. They are ordinary engineering discipline pointed at writing: state the requirement, inspect the output, record the exceptions.
Somebody had printed this out and stuck it on the wall of my school computer lab. I memorised it while learning to code in Pascal, and I have thought of it most weeks since.
I really hate this damn machine,
I wish that they would sell it.
It never does what I want it to,
But only what I tell it.Author unknown
A good deal of my working life at present is spent on a related problem: how often a business or a process actually looks at its own numbers. The failure mode there has the same shape as this one. Data sampled too slowly does not go quiet. It produces a confident, stable, entirely plausible picture of something that is not happening.
The mechanisms are not identical. One is about the interval between observations, the other about which observations reach you at all. But they fail the same way from the inside, and that is the part worth carrying: a systematically incomplete sample does not feel incomplete. It feels like knowledge.
Which is why "you can always spot AI content" is said with such confidence. The evidence really is overwhelming. It is simply evidence for a different proposition than the one being advanced.
Under the rules above. Several drafts, one argument I had to reconsider, and a check on the Wald account before publishing. If you reached this line without noticing, that is the whole point, and you are entitled to hold it against me.
Start a Conversation