Turnitin Tested Its Own Detector for Bias Against English Learners. The Result Has Two Halves.

For submissions over 300 words, Turnitin reports the gap between L2 and L1 writers is small and not statistically significant. For submissions between 150 and 300 words, the same evaluation found a wider gap, significantly above its own 1% target. And Turnitin states a limitation about its sample that matters if you write at university length.

HumanPen Team

· 6 min read

The short answer, both halves of it

Turnitin ran its own evaluation on this question and published the result. The answer has two halves, and they point in opposite directions.

For submissions of 300 words or more, it found the gap in false positive rate between English Language Learners and native English writers was small and not statistically significant. For submissions between 150 and 300 words, the same evaluation found a wider gap, significantly above its own 1% target.

So whether the reassuring finding covers you depends on how long the thing you submitted was. There is also a limitation Turnitin states about its own sample that matters if you are writing at university level, and it is further down.

What was tested

Turnitin pulled up to 2,000 texts for each combination of two variables: L2 English (English Language Learner) writing versus L1 English (native English) writing, and "too short" (between 150 and 300 words) versus "long enough" (300 words or more). The datasets were not part of their detector's training set.

The result that goes in the headline

For documents that meet the 300-word minimum, Turnitin reports that the difference in false positive rate between L2 and L1 writers is small and not statistically significant:

"For documents meeting our minimum 300 word count requirement, the difference in the false positive rate (FPR) between L2 English writers (ELL writers) and L1 English writers (native English writers) is small, and not statistically significant, illustrating that our detector is not biased between these two groups of writers. Additionally, although the FPR for each group is slightly higher than our overall target of 0.01 (1%), neither group's FPR is significantly different from this target."

"Not statistically significant" means the difference they found could be noise. It does not mean the difference is zero, and it does not mean no bias exists. And both groups' false positive rate was above the 1% target, just not by a margin the test treated as meaningful. "Slightly higher" is not "at or below."

The result that doesn't go in the headline

For documents between 150 and 300 words, the same evaluation found the opposite direction. The gap between L2 and L1 writers was wider, and the false positive rate was significantly above the 1% target:

"Conversely, for documents that are "too short," we found there is a greater difference in FPR between these groups of writers and it's significantly greater than our target of 0.01 FPR. This again corroborates our earlier finding that AI writing detectors require longer samples to correctly identify AI-generated content and to keep false positives to a minimum."

The same study, on the same detector, produced one result for longer texts and a different result for shorter ones. The blog title carries the first. The second is in the body.

What was actually in the sample

Turnitin states:

"Most of the L2 English documents are short persuasive or informative writing tasks, most similar to the Automated Scoring and Assessment Prize (ASAP) secondary-level essay corpus. A comparable collection of publicly available full-length university level documents from L2 and L1 writers was not available for this evaluation."

The L2 writing in this evaluation is mostly secondary-school-level essays. If you're writing a university paper, a literature review, a lab report, a thesis chapter, that's not what was tested. Turnitin says they couldn't find a publicly available collection of university-level L2 and L1 writing to run this comparison on. The "not statistically significant" finding applies to short persuasive and informative essays at a secondary level. Whether it holds for the kind of document you're actually submitting is a question this evaluation can't answer.

The 300-word line

Turnitin traces the 300-word minimum directly to this research. After evaluating 800,000 documents and seeing that short submissions pushed the false positive rate up, they moved to require 300 words:

"Our evaluation of 800,000 documents—mentioned earlier in this blog—shows a slightly higher-than-comfortable false positive rate on short submissions (fewer than 300 words). There may not be enough signal/markers in such short samples to identify the telltale distributional differences in AI writing. In order to ensure we keep our false positive rate below 1%, we moved quickly to update the minimum submission length to 300 words for our AI writing detection capability to process a submission."

The minimum exists because short texts caused problems. The reassuring result only applies above that minimum. Below it, the same evaluation found the gap you'd want to know about.

What about the Liang study?

You may have seen headlines tied to a study by Liang. Turnitin addresses it directly:

"Liang's claim is grounded on a small collection of works of just 91 Test of English as a Foreign Language (TOEFL) practice essays, all of which are less than 150 words long. Turnitin's AI writing detection was not included in this evaluation, perhaps in part because we won't make predictions about documents that are so short."

What to do with this

If your submission is over 300 words, Turnitin's own data says the gap between L2 and L1 false positive rates is small and not statistically significant. That's not a guarantee you won't be flagged. It means the evaluation didn't find strong evidence that L2 writers are systematically worse off, at least not in the sample they had.

Here's what you can actually use:

Check your word count before submitting. If a section under 300 words gets flagged, the detector's own evaluation says short texts produce a wider gap between L2 and L1 writers and a higher false positive rate overall. You can ask your instructor to run the check on the full document rather than a short excerpt.

Ask what flagged you and at what threshold. If Turnitin was used, the 300-word minimum and the "slightly higher than 1%" detail are specific things you can raise with whoever is reviewing the flag.

Keep your drafts, notes, and sources. If you're flagged and you think your L2 writing played a part, showing your revision history and source materials is what resolves the conversation. Arguing about statistical significance with your professor probably won't.

KEEP READING