Why non-native English writers get flagged by AI detectors more often

Seven detectors called 61% of human-written TOEFL essays AI. The reason is not cheating - it is that detectors measure vocabulary range, and a controlled experiment in the same paper proves it.

HumanPen Team

· 9 min read

The study: 61.22% of human essays called AI

In 2023 a Stanford group - Liang, Yuksekgonul, Mao, Wu and Zou - ran seven widely used GPT detectors over two sets of essays. The first was 91 TOEFL essays taken from a Chinese educational forum, all written by students. The second was 88 essays by US eighth graders, from the Hewlett Foundation's ASAP dataset. No AI was involved in either set. The work was published in Patterns, a Cell Press journal.

On the US eighth-grade essays the detectors were near-perfect. On the TOEFL essays the average false positive rate was 61.22%. All seven detectors unanimously called 18 of the 91 essays (19.78%) AI-written, and 89 of the 91 (97.80%) were flagged by at least one.

Read the control again, because it carries the argument. The same seven tools, the same day, were accurate on native-speaker writing and wrong most of the time on second-language writing. They were not broken. They were unreliable on one population and not the other.

The essays came from a Chinese educational forum. If you are writing English as a second language, this study is not about someone else.

The finding has been contested - a detector vendor argues 91 essays from a student forum is too small a sample to carry the rate. That is a fair thing to say about any single study. But institutions have already acted on the direction. Vanderbilt University disabled Turnitin's AI detector in August 2023, and its arithmetic is worth repeating: Turnitin advertised a 1% false positive rate, Vanderbilt received roughly 75,000 papers in 2022, and one percent of that is about 750 students wrongly accused in a single year. The university also said it had no visibility into how the score was produced.

What detectors actually measure

Most detectors estimate two statistics, and neither of them is authorship.

Perplexity measures how surprising your word choices are to a language model: if the next word is the one the model would have predicted, perplexity is low. Burstiness measures how much sentence length and structure vary across a passage; even, regular sentences score low.

Text that is low on both looks machine-like to a detector. That is the entire test.

A detector never sees intent. It sees two statistics - and careful, textbook-correct English sits closer to model output than a native speaker's looser prose does.

So this is not a defect a better model will fix. The measurement does what it was designed to do. It simply does not measure the thing a reader assumes it measures.

The experiment that proves it is vocabulary, not authorship

The paper does not stop at the correlation. It runs the causal test in both directions, and this is the part worth knowing.

Direction one. The researchers took the TOEFL essays that had been unanimously flagged and asked GPT-4 to enrich the word choices. Perplexity rose significantly (P = 9.36E-05), and the number unanimously detected fell from 18 essays to 1 of 91 (1.10%). Same authors, same arguments, same essays - only the vocabulary changed.

Direction two, which is the stronger one. They took the US eighth-grade essays, which detectors had judged almost perfectly, and prompted ChatGPT to "simplify word choices as if written by a non-native speaker." Misclassification jumped from an average of 5.19% to 56.65%.

Native-written text, made lexically simpler, becomes AI to a detector. That is not a correlation between demographics and scores. It is the measurement responding to word variety, with authorship held constant.

The paper's own conclusion is that detectors "may unintentionally penalize writers with constrained linguistic expressions," and it cautions against deploying them in evaluative or educational settings at all.

A related finding: one prompt defeats them

The same paper reports something that complicates any use of these scores as evidence. The team generated essays with ChatGPT-3.5 - detectors caught them reliably. Then they applied a single self-edit prompt, "elevate the provided text by employing literary language," and detection fell from 100% to 13%.

This is not offered here as a technique. It is offered as a fact about the instrument: a measure that a one-line prompt moves by 87 percentage points is a weak basis for an accusation, and it is also why nobody can honestly promise you a particular detector outcome. We do not.

What to do if you are flagged

None of this is advice about writing around a detector. It is what to do when a statistic is being read as evidence about you.

  • Keep your drafts, starting now. Version history in Word or Google Docs, outline files, reading notes. Provenance is the one thing a detector score cannot argue with, and it only exists if you kept it before you needed it.
  • Ask for the full report, not the headline percentage. Section-level scores usually show the flag concentrated in passages that are supposed to read formulaically - methods, definitions, background. That distribution is itself an argument.
  • Ask how the number is being used. Many institutions treat a score as a reason to look, not as a finding; Vanderbilt's guidance says so in writing. Asking what your own institution's policy is, is a reasonable question.
  • Write with deliberate variation. Mixing long and short sentences, and choosing the precise word over the safe general one, is good writing advice whether or not anyone is running a detector. It happens to raise both statistics.

If you already have a report and want to see exactly which passages it flagged rather than guess, importing the report does that: the flagged passages are located, and the rest of the document is left untouched.

KEEP READING