On Working With AI

The Content You
Did Not Notice

Why "you can always spot AI writing" tells you almost nothing about how much AI writing you have read.

The Observation

A Claim You Have
Probably Made

"It is so easy to spot AI generated content. It always sounds the same."

I hear some version of this most weeks. It came up again recently, typed into a chat window during a presentation about AI, by somebody entirely confident about it.

And they are right. Anyone who reads a reasonable amount has built a working instinct for it. The particular cadence. The list of three. The paragraph that restates the previous paragraph in slightly different words. The closing sentence that gestures at significance without adding anything. The em dashes and the en dashes. Once you have seen it a few times you cannot unsee it.

So the observation is sound. It is the conclusion drawn from it that does not follow.

"You are not judging a sample of AI writing. You are judging a sample of AI writing that failed."

Every piece you have ever correctly identified has one property in common, and it is the property that put it in front of you.

I am probably worse at this than I think, in both directions. I have flagged writing as machine made that almost certainly was not, and I never find out about those either. Neither kind of error reports back.

EVERYTHING YOU READ did it carry the tells? WHAT YOU NOTICED

The pale ones are not absent from the world. They are absent from the sample you formed the judgement on, and they are absent for exactly the reason you would want to know about.

The Short VersionYour detection rate tells you nothing about the base rate. You can be right about every instance you have caught and still have no idea what proportion of what you read was machine assisted.
The Mechanism

The Sample Has
Already Been Filtered

There is a well known version of this error, and it is worth walking through, because everybody involved in it was clever and looking carefully at real data.

The Aeroplanes

Where The Bullet Holes Were

During the Second World War, analysts examined bombers returning from missions over Europe and mapped where they were taking damage. A clear pattern emerged: wings, fuselage, and the area around the tail gunner. The sensible response was to add armour where the bullet holes were.

Abraham Wald, a statistician working for the Statistical Research Group in New York, pointed out that they were examining the wrong aeroplanes. The map showed where a bomber could be hit and still get home. The aircraft hit in the engines were not in the sample. They had not come back, and neither had their crews.

The armour went on the engines.

NO DAMAGE RECORDED HERE

Damage recorded on the aircraft that came back. The engines are not undamaged. They are unrepresented.

The Writing

Where The Tells Were

Now apply the same reading to your instinct about AI writing, or about any other kind of content if that is your domain.

Every piece you have correctly identified failed. It carried the tells, nobody removed them, and it reached you in that state. That is not incidental to your sample. It is the entry requirement.

The pieces that were carefully specified, reviewed, corrected and rewritten before publication did not announce themselves either. If the work was done properly you read them and thought nothing at all, which is the outcome the work was for. They are not in your sample and they cannot be, because the thing that would have placed them there is precisely the thing that was removed.

The error is not a failure of intelligence in either case. It is that absence does not announce itself. A sample that has been filtered feels exactly like a sample that has not, from the inside.

What It Actually Measures

A Different Property
Than You Think

There is a version of this worth saying plainly, because it changes what the complaint is actually about.

When somebody says AI writing is obvious and bad, the observation underneath is accurate. What they encountered was obvious and bad. But the property being detected is not "written with AI." It is "published without review."

Those two things correlate strongly at the moment, which is why the instinct works as well as it does. They are not the same property, and the correlation is a fact about current habits rather than about the tool.

A first draft from a language model, shipped unread, has a signature. So does a first draft from a person, shipped unread. The difference is that nobody has built a folk theory about the second one, because it has been happening for as long as there has been writing, and we have other words for it.

Why This Is Not Just A Logic PuzzleThe belief produces a real decision. People conclude that AI cannot write acceptably, when what they have observed is that unreviewed work reads as unreviewed. Those lead to very different choices about how to use it.
A Fair ObjectionSome readers will say the tells are inherent to the models rather than to the reviewing. That is a reasonable position and it is testable: it predicts that careful review cannot remove them. My own experience is that it can, but I hold it as experience rather than proof.
What Closes The Gap

The Work Sits
With The Human

I use AI extensively. Most days, across engineering work, client documents, code and writing. I am not a neutral party here and there is no point pretending otherwise.

What I have found is that the difference between output carrying the tells and output not carrying them is almost never the model. It is whether somebody specified what they actually wanted, and then checked whether they got it.

That still relies, for now, on a person being able to carry the specification. It is a partnership rather than a handover, and the responsibility sits on the human side of it.

Spell check has been on nearly every machine most of us have used for thirty years and it has never been perfect. Some of that is the tool. A good deal of it is somebody leaving the dictionary set to the wrong variant of English and then blaming the red underline.

None of the rules alongside are sophisticated. They are ordinary engineering discipline pointed at writing: state the requirement, inspect the output, record the exceptions.

PunctuationNo em dashes in anything public facing. They are a reliable tell, they render inconsistently across devices, and the alternatives force clearer sentence structure. Colons, full stops, commas, parentheses, semicolons. This page contains none.
Review Before HandoverAnything I am going to deploy gets a review pass first, with a written note of what was checked, what was found, and what remains uncertain. That note is the decision point before I spend my own time on it.
Verification Of Present Tense ClaimsAny statement about the current state of a system I run is either verified against something live or explicitly flagged as unverified and document sourced. This rule exists because I got it wrong once.
Instructed DisagreementThe standing instruction is to push back directly when something I propose has a problem, rather than agreeing and proceeding. An assistant that agrees with everything is simply a faster route to something confidently wrong. Several lines on this page changed because one of us argued with the other.

Somebody had printed this out and stuck it on the wall of my school computer lab. I memorised it while learning to code in Pascal, and I have thought of it most weeks since.

I really hate this damn machine,
I wish that they would sell it.
It never does what I want it to,
But only what I tell it.Author unknown
The Same Shape

Incomplete Samples
Do Not Feel Incomplete

A good deal of my working life at present is spent on a related problem: how often a business or a process actually looks at its own numbers. The failure mode there has the same shape as this one. Data sampled too slowly does not go quiet. It produces a confident, stable, entirely plausible picture of something that is not happening.

The mechanisms are not identical. One is about the interval between observations, the other about which observations reach you at all. But they fail the same way from the inside, and that is the part worth carrying: a systematically incomplete sample does not feel incomplete. It feels like knowledge.

Which is why "you can always spot AI content" is said with such confidence. The evidence really is overwhelming. It is simply evidence for a different proposition than the one being advanced.

More on the sampling problem

The Disclosure

This Page Was Written With AI

Under the rules above. Several drafts, one argument I had to reconsider, and a check on the Wald account before publishing. If you reached this line without noticing, that is the whole point, and you are entitled to hold it against me.

Start a Conversation