Why non-native English writers get flagged by AI detectors more often
Seven detectors misclassified 61% of a sample of human-written TOEFL essays. The study does not represent every multilingual writer, but its controlled edits show why limited lexical range can move a detector score without changing who wrote the argument.
HumanPen Team
· 8 min read
The study: 61.22% of human essays called AI
In 2023 a Stanford group - Liang, Yuksekgonul, Mao, Wu and Zou - ran seven widely used GPT detectors over two sets of essays. The first was 91 TOEFL essays taken from a Chinese educational forum, all written by students. The second was 88 essays by US eighth graders, from the Hewlett Foundation's ASAP dataset. No AI was involved in either set. The work was published in Patterns, a Cell Press journal.
On the US eighth-grade essays the detectors were near-perfect. On the TOEFL essays the average false positive rate was 61.22%. All seven detectors unanimously called 18 of the 91 essays (19.78%) AI-written, and 89 of the 91 (97.80%) were flagged by at least one.
The contrast is the important result. The same seven tools performed much better on the US eighth-grade set than on the TOEFL set. That does not establish an error rate for every multilingual writer or every detector used today. It does show that an aggregate accuracy claim can hide a much worse result for a particular writing population.
Read 61.22% as a result for this sample and these seven 2023 detectors, not as your personal probability of being flagged. The study is a warning about unequal error, not a universal forecast.
The finding has been contested. A detector vendor argued that 91 essays from one student forum were too few and too narrowly sourced to support a general rate. Those are real limitations: the sample is a convenience sample, the comparison groups differ in age as well as language background, and detector products have changed since 2023. The defensible conclusion is therefore narrower than the headline: the study identified a large disparity that institutions should not assume away.
Institutions have reached the same caution from a risk perspective. Vanderbilt University disabled Turnitin's AI detector in August 2023. It noted that even a claimed 1% document-level false-positive rate would create a serious review burden across roughly 75,000 annual submissions. That arithmetic is an illustration, not a prediction that exactly 750 students would be accused: not every paper is eligible, the base rate is unknown, and a flag need not become a charge. The point is that small percentages still require a fair process at institutional scale.
What would strengthen the evidence? Larger, pre-registered studies should sample writers across first languages, proficiency levels, disciplines and assignment types; compare age-matched groups; publish results per detector version; and report both false positives and false negatives. Aggregate accuracy alone is not enough. Institutions considering deployment should ask whether the vendor has evaluated the actual student population and whether subgroup results are available, rather than importing a headline rate from a different corpus.
For an individual report, neither the Stanford number nor a vendor benchmark settles the case. The relevant evidence is local: what text qualified, where the highlights fall, which model version ran, what the assignment allowed and how the student's drafts developed. Population research tells a reviewer which failure mode to take seriously; it does not replace review of the particular work.
What detectors actually measure
Early public detector explainers focused on two statistical signals: perplexity and burstiness. Modern commercial systems may combine many learned signals, and their internals are usually not disclosed, so these two terms are best treated as a way to understand the problem rather than a complete description of every current product.
Perplexity measures how surprising your word choices are to a language model: if the next word is the one the model would have predicted, perplexity is low. Burstiness measures how much sentence length and structure vary across a passage; even, regular sentences score low.
In a detector that uses these signals, low perplexity and low variation can contribute to a machine-like classification. They are not a universal test, and even the same passage can have different perplexity under different reference models. A classifier still has to combine its signals and choose a threshold at which to label the result.
A detector does not observe authorship or intent. It observes features of the submitted text and compares them with patterns learned from training data.
That distinction matters for second-language writing. A smaller working vocabulary, repeated sentence frames and strict adherence to classroom models can resemble features found in generated text. A better classifier may reduce that error, but no architecture turns a textual resemblance into direct observation of who wrote the passage.
The threshold matters too. Lowering it catches more generated text but also admits more human text; raising it does the reverse. An advertised accuracy number describes one dataset at one operating point. It does not guarantee equal performance across proficiency levels, disciplines, document lengths or later model versions.
The experiment showing that lexical range moves the result
The paper does not stop at comparing two groups. It manipulates the language in both directions while keeping the source essays and their arguments fixed. That makes the proposed mechanism more persuasive, although it still does not isolate vocabulary from every other stylistic change made by the model.
Direction one. The researchers took the TOEFL essays that had been unanimously flagged and asked GPT-4 to enrich the word choices. Perplexity rose significantly (P = 9.36E-05), and the number unanimously detected fell from 18 essays to 1 of 91 (1.10%). Same authors, same arguments, same essays - only the vocabulary changed.
Direction two, which is the stronger one. They took the US eighth-grade essays, which detectors had judged almost perfectly, and prompted ChatGPT to "simplify word choices as if written by a non-native speaker." Misclassification jumped from an average of 5.19% to 56.65%.
Authorship of the underlying essays did not change, but the classifications did. The intervention therefore supports a causal role for surface language - especially lexical variety - rather than a demographic label somehow being visible to the detector. Because GPT-4 and ChatGPT may also change syntax, cadence and phrasing, it would be too strong to say the experiment proves vocabulary is the only cause.
The paper's own conclusion is that detectors "may unintentionally penalize writers with constrained linguistic expressions," and it cautions against deploying them in evaluative or educational settings at all.
The result also explains why asking a multilingual writer to sound more formal can backfire. Academic advice often rewards standard transitions, conservative vocabulary and uniform sentence structure. Those choices may be appropriate for clarity, yet they reduce exactly the variation that older statistical detectors treated as human evidence. The remedy is not to force decorative vocabulary into a paper; it is to avoid treating the detector as an authorship test.
A related finding: one prompt defeats them
The same paper reports something that complicates any use of these scores as evidence. The team generated essays with ChatGPT-3.5 - detectors caught them reliably. Then they applied a single self-edit prompt, "elevate the provided text by employing literary language," and detection fell from 100% to 13%.
This is not offered here as a technique. It is offered as a fact about the instrument: a measure that a one-line prompt moves by 87 percentage points is a weak basis for an accusation, and it is also why nobody can honestly promise you a particular detector outcome. We do not.
There are two limits on that result. First, it describes the seven detectors and model outputs tested in 2023, not every product in 2026. Second, robustness to evasion and fairness to human writers are different properties: a detector can become harder to fool and still retain unequal false positives. Any current evaluation should test both, on the populations and writing genres where the score will actually be used.
A detector can be useful for triage without being reliable enough for adjudication. The higher the consequence, the more independent evidence and procedural review are needed.
What to do if you are flagged
None of this is advice about writing around a detector. It is what to do when a statistic is being read as evidence about you.
- Keep your drafts, starting now. Preserve version history in Word or Google Docs, outline files, reading notes and supervisor feedback. These artefacts show how the work developed over time; they are stronger evidence than a competing detector score.
- Ask for the full report, not the headline percentage. Passage highlights can show whether flags cluster in methods, definitions or background - sections expected to use conventional language. Do not describe them as section-level probabilities unless the report actually provides those.
- Ask how the number is being used. Request the relevant course or institutional policy, the detector and version, and the evidence considered beyond the score. Many institutions treat a flag as a reason to look, not as a finding.
- Edit for meaning and ownership. Replace wording you would not naturally use, check every claim and citation, and make sure you can explain the argument. Do not distort clear prose merely to manufacture sentence-length variation.
If a formal process has started, keep the submitted file unchanged and work from a copy. Write down a simple timeline: when you researched, drafted, received feedback and revised. Ask for enough time to review the evidence, and bring an adviser or student representative where the rules permit it. The separate guide on preparing a response to a false flag covers that process in detail.
If you are an educator, do not wait for a student to raise the bias question. Define in advance who reviews a flag, what corroboration is required, how language background and disability are considered, and how the student receives the underlying report. A tool can be optional triage only if the human process is capable of disagreeing with it.
KEEP READING