Will an AI Detector Flag You If English Isn't Your First Language? What GPTZero's 0-of-914 Test Shows
If English is your second language, you've probably read that AI detectors flag people like you more often. On 5 October 2026 two new sets of numbers came out on the same day, one from GPTZero about its own model and one from university researchers about Pangram. Both say almost no English learners' essays were called AI. Here is what each test used, what "zero" can and can't prove, and why the 2023 study that started the worry still stands next to them.
HumanPen Team
· 12 min read
Will an AI detector flag your essay if English isn't your first language?
For two detectors, on one kind of essay, the newest numbers are very low. GPTZero says its current model, GPTZero 4o, labeled as human-written all 914 essays by English learners in a test set it hadn't been trained on. A preprint posted the same day by researchers at the University of Maryland, Google DeepMind and Simon Fraser University ran Pangram 4 on 3,909 essays by English learners and flagged 0.0% of them. Both sets were written by US secondary-school students, and neither study tested Turnitin.
Read the zero with its fine print. GPTZero ran and published its own test, and 0 out of 914 is still consistent with a real rate of up to about 1 in 300. It also doesn't cancel the 2023 Stanford study that started this worry. There, the version of GPTZero online in March 2023 flagged 52% of 91 TOEFL essays, every one of them under 150 words. The model, the essays and their length all differ, so the two results don't contradict each other. Neither one is a promise about your essay.
What GPTZero tested
The post is "Biased Against ESL and Students with Disabilities? We Put GPTZero 4o to the Test", by Anna Gao, on GPTZero's blog on 5 October 2026. GPTZero announced 4o on 24 September and says it became the default model for all users on 20 September.
The essays come from PERSUADE 2.0, a research corpus of more than 25,000 argumentative essays by US students in grades 6 to 12, built to improve automated writing feedback. Each essay carries information about its writer, including whether the student was classed as an English language learner and whether they had an IEP or a Section 504 plan (US school support plans for students with disabilities). Version 2.0 adds scores to the essays first released as PERSUADE 1.0, which only kept essays of at least 150 words in which at least 75% of the words were correctly spelled. They average 402 words. PERSUADE 1.0 was published in a journal in October 2022, before OpenAI released ChatGPT on 30 November 2022, so none of these essays could have been written with ChatGPT.
GPTZero scored essays from what it calls "PERSUADE2.0's official held-out test split, which GPTZero 4o was not trained on." That included 914 essays by English learners and 1,172 by students with an IEP or 504 plan, 147 of them in both groups. GPTZero 4o labeled every one of them human-written. Its charts put all 914 English learners' essays in the single bar at the top of its "Human probability" scale, and for the 147 students in both groups the lowest score was 0.997.
Two things the post doesn't say. It doesn't give the score at which an essay counts as "human-written". And it doesn't say whether the rest of PERSUADE went into training 4o. That training portion has been public since a Kaggle data-science competition, the Feedback Prize 2021, and it comes from the same corpus of argumentative essays by US 6th to 12th graders. So the held-out split is a fair check on essays like those, and a weaker one on a university lab report.
It's also GPTZero's own test of GPTZero, on a dataset it chose, published on its own blog, with no link to code or to the scored essays. That doesn't make the numbers wrong. It means you're reading the company's account of its product, and GPTZero itself calls the result "a great milestone; however, it is not a finish line."
What 0 out of 914 can and can't tell you
Zero out of 914 is a count, and it doesn't pin down a rate. With 914 essays and no false positives, the true rate could still be as high as about 0.33%, roughly 1 in 300. That's the usual 95% upper bound when something happens zero times in 914 tries. Low, but not zero.
The second limit is less obvious. GPTZero's 4o launch post reported "only 1 false positive out of 10,402 essays" in the PERSUADE test set. That test split has 10,402 essays in total, so it's the same set the 914 came from. By our arithmetic, if English learners were flagged exactly as often as everyone else, you'd expect about 0.09 false positives among 914 essays, which means you'd see zero almost every time. If they were flagged ten times as often, you'd still see zero at least four times in ten.
So 0 out of 914 can't show that English learners and other students are treated the same. A test this size couldn't tell those situations apart. What it does show is that on these essays the false-positive rate was very low for everyone, English learners included.
The independent test: Pangram 4 on 3,909 essays
The second set of numbers is in an appendix of "IdeaLens: Detecting AI Ideas in Long-form Writing", a preprint by Rishanth Rajendhran and seven co-authors, posted to arXiv on 5 October 2026. The paper is mainly about a research detector of the authors' own. Pangram 4, a commercial detector, is one of the tools they compared it with across 19 sets of ordinary human writing. Pangram didn't run the test. The authors do thank it for a research credit award, and the paper hasn't been peer reviewed.
One of those 19 sets is ELLIPSE, a corpus of about 6,500 essays by English learners in grades 8 to 12. Each was written in 25 to 30 minutes on a computer during annual state-wide tests in the US. Trained raters scored every essay for English proficiency from 1 to 5, where 1 means limited ability and 5 means "native-like facility". The paper used 3,909 of them. When we checked, that number and its breakdown by grade and by score match the corpus's public training file to within one or two essays (the file has 3,911). About 70% of the 3,909 are under 500 words.
The authors count a text as AI when Pangram's AI share plus half its AI-assisted share comes to at least half. By that rule, Pangram flagged 0.0% of the 3,909. The figure is rounded to one decimal place, so it allows for one essay at most. In the authors' published results file, it was 0.0% in every proficiency band as well, though the bands are uneven: the lowest holds 8 essays and the highest 28, while bands 2, 3 and 4 hold more than a thousand each.
Here a base-rate check actually works, because the same paper reports Pangram on ordinary human writing in general. Averaged over all 19 evaluations, Pangram flagged 0.3% of human-written documents. That average comes mostly from two sets, a mixed benchmark of human texts and a set of expert reviews of research proposals. Fifteen of the 19, ELLIPSE among them, came in at 0.0%. The other sets range from peer reviews to web pages, so this compares the learners with Pangram's behaviour on human writing in general, not with native-speaker classmates. That is the closest either new test comes to answering whether English learners get flagged more than other writers. For Pangram, on these essays, they didn't.
Four tests of English learners' writing, side by side
| Test | Detector and when | Essays by English learners | Flagged as AI |
|---|---|---|---|
| Liang and colleagues, Stanford (published 2023) | GPTZero as accessed on 15 March 2023, one of seven detectors | 91 TOEFL essays from a Chinese education forum, 62 to 148 words | 52% (61.22% averaged over all seven detectors) |
| Turnitin's own evaluation (October 2023) | Turnitin, the version available to customers at the time | 2,221 texts of 300 words or more, from university-level learners | 30 (1.4%), against 15 of 1,155 (1.3%) for native writers |
| GPTZero's own test (October 2026) | GPTZero 4o | 914 essays, US grades 6 to 12 | 0 |
| IdeaLens preprint (October 2026) | Pangram 4 | 3,909 essays, US grades 8 to 12 | 0.0% (one essay at most) |
This is not a ranking. Every row is a different detector or a different version, each with its own definition of "flagged as AI", run on different writers. The word counts in the first row are ours, from the essays the Stanford authors published with their code. We covered that study in Why non-native English writers get flagged by AI detectors more often and Turnitin's evaluation in Turnitin Tested Its Own Detector for Bias Against English Learners. The Turnitin row describes the 2023 model. Turnitin has changed its detector since then.
Does this mean the bias problem is gone?
Not on this evidence. The 2026 numbers are real, but they cover a narrower situation than "English learners" in general.
- Length. Every essay in the Stanford set was under 150 words. Turnitin's 2023 test found the gap between learners and native writers widened for texts of 150 to 300 words. The two new sets mostly sit well above that: PERSUADE essays average 402 words, and 30% of the ELLIPSE essays run to 500 words or more. The very short writing behind the alarming 2023 figure isn't what the 2026 tests measured.
- Level. Both new sets are secondary-school writing. Turnitin's 2023 learner texts did come from university-level learners (its dataset table labels both learner corpora "higher ed"), but by its own description they were mostly short persuasive or informative tasks, and almost all of its native writers in the 300-word group were US secondary students. It also wrote: "A comparable collection of publicly available full-length university level documents from L2 and L1 writers was not available for this evaluation."
- The weakest writers are thin on the ground. PERSUADE dropped essays with too many misspellings, and ELLIPSE's lowest proficiency band has 8 essays.
- Turnitin. Its AI detection FAQ says that in July 2026 it folded a multi-model ensemble into one model, and its answer to the bias question describes how training data was chosen, with no figures. In the FAQ's text, "English learner", "ESL", "L2" and "native" each appear 0 times, and the four mentions of "second language" come with no numbers.
- Versions. GPTZero 4o has only been the default since 20 September. The paper doesn't say exactly when Pangram scored the essays. Both detectors will be updated.
The fair summary is narrower than the headlines: in autumn 2026, two detectors, on essays of a few hundred words by US secondary-school learners, almost never called them AI.
If English isn't your first language and you're worried about a flag
- Find out which detector was used, and when. These numbers describe GPTZero 4o and Pangram 4. If your school uses Turnitin, or the check was run on GPTZero before 20 September 2026, they don't carry over.
- Ask for the full report, not just a label or a percentage. You want to see which passages were marked and what the score was.
- If the flagged piece is short, mention it. The older studies found their problems in short writing, and Turnitin's own report won't show an AI percentage for fewer than 300 words of prose. A flag on a short discussion post rests on much less text than one on a full essay.
- Hold on to your drafts. Earlier versions, notes, the sources you read and version history are evidence about your essay. A study's rate across other people's essays isn't. How to keep version history in Word, Google Docs and Overleaf explains which setups actually keep it.
- Quote the studies accurately if you bring them up. If GPTZero flagged you, it's fair to say GPTZero's own published test labeled all 914 English learners' essays human. Offer it as a reason to look at your drafts, not as proof. A test of 914 school essays can't settle what happened with yours.
Frequently asked questions
Does this mean Turnitin won't flag my essay? These studies can't tell you. Turnitin's own 2023 evaluation found 1.4% of English learners' texts of 300 words or more flagged, close to the 1.3% for native writers, on a model it has since replaced, and its current FAQ gives no figures for English learners.
I'm at university. Do these results apply to me? Only loosely. Both 2026 sets are essays by US school students: classroom assignments in PERSUADE, and 25-to-30-minute test essays in ELLIPSE. A university literature review is longer, more technical and written over weeks. The closest data is Turnitin's 2023 evaluation, whose learner texts came from university-level learners, though mostly from short tasks: 1.4% of those of 300 words or more were flagged, on a model Turnitin has since replaced.
Should I make my English simpler, or fancier, to avoid a flag? Nothing in these results points that way. In the Pangram test, essays at proficiency levels 2, 3 and 4, more than a thousand at each, all came out at 0.0%, and the same held across the grammar and vocabulary score bands in the authors' results file. That's one detector, but it's no reason to write in a voice that isn't yours.
What about students with disabilities? The same GPTZero post reports that 4o labeled all 1,172 essays by students with an IEP or 504 plan as human-written. The same caveats apply: a vendor's own test, school essays, and a count of zero that still allows a small real rate.
Sources
- Anna Gao, Biased Against ESL and Students with Disabilities? We Put GPTZero 4o to the Test, GPTZero, 5 October 2026 (text and the five images on the page); GPTZero Team, Introducing GPTZero 4o, 24 September 2026; both read 7 October 2026.
- Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz Bölöni-Turgut, Marzena Karpinska, John Wieting and Mohit Iyyer, IdeaLens: Detecting AI Ideas in Long-form Writing, arXiv:2610.06778v1, 5 October 2026 (Appendix E.3, Section G.1, Table 17), and the authors' ELLIPSE results file; read 7 October 2026.
- Scott Crossley and colleagues, A large-scale corpus for assessing written argumentation: PERSUADE 2.0, Assessing Writing 61 (2024), preprint; The persuasive essays for rating, selecting, and understanding argumentative and discourse elements (PERSUADE) corpus 1.0, Assessing Writing 54 (2022); read 7 October 2026.
- Scott Crossley and colleagues, The English Language Learner Insight, Proficiency and Skills Evaluation (ELLIPSE) Corpus, International Journal of Learner Corpus Research 9(2), 2023, preprint and corpus files; read 7 October 2026.
- OpenAI, ChatGPT: Optimizing Language Models for Dialogue, 30 November 2022 (Internet Archive copy of 2 December 2022); read 7 October 2026.
- Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou, GPT detectors are biased against non-native English writers, arXiv:2304.02819v3 (Figure 1a and Materials and Methods; published in Patterns, 2023), and the authors' data repository; read 7 October 2026.
- David Adamson, New research: Turnitin's AI detector shows no statistically significant bias against English Language Learners, Turnitin, 26 October 2023 (text and the results table image); Turnitin, AI writing detection capabilities FAQs and Using the AI Writing Report; read 7 October 2026.
KEEP READING