Does Turnitin Flag Non-Native English Writing? A Study Published in 2026

If English is not your first language, you have probably heard that AI detectors are biased against you. A study published in 2026 by researchers at Sultan Qaboos University in Oman tested that with Turnitin and Originality, using coursework written by its own foundation-year students before ChatGPT existed. Here is what it found for student writing, for mixed human and AI text, and for science compared with humanities, and what that means for a score you have been given.

HumanPen Team

· 5 min read

Does Turnitin flag non-native English writing?

Sometimes, in this study, and at about the same rate as BBC news writing. Of 48 essays by English-as-a-foreign-language (EFL) students, all written before August 2022, Turnitin scored 4 above 20%, and so did Originality. On 48 professionally written BBC news texts, Turnitin scored 3 above 20% and Originality none. The authors found no significant difference for Turnitin between the two groups, and a borderline one for Originality. Both detectors did poorly on texts that were half human and half AI.

The study is "Evaluating the accuracy and reliability of AI content detectors in academic contexts", by Mohammad Hadra, Karleen Cambridge and Mostefa Mesbah of Sultan Qaboos University, published open access in the International Journal for Educational Integrity on 2 February 2026. The authors declare no competing interests.

What the researchers tested

They built 192 texts in four groups of 48:

  • EFL student writing. Final coursework from the university's Foundation Program English for Sciences courses, all submitted before August 2022, before tools like ChatGPT were available to students.
  • Professional writing. Texts from XSum, a research dataset of BBC news articles.
  • AI-generated. Texts produced by two large language models on essay tasks like the students'.
  • Hybrid. About half EFL student writing and half AI-generated text, joined into one piece.

The texts were 300 to 1,100 words long, and 135 of the 192 were on science topics. The detectors were run between January and May 2025. Each detector's score was sorted into three bands: 0 to 20% counted as Human, 21 to 79% as Hybrid, and 80% or more as AI.

What it found for human writing

Human-written textsTurnitin: scored 20% or lessOriginality: scored 20% or less
EFL student essays (48)44 (91.6%)44 (91.6%)
BBC news texts (48)45 (93.8%)48 (100%)

For Turnitin, the gap between student essays and news writing was not statistically significant (p = 0.50). The authors write that this "suggests that Turnitin did not exhibit measurable bias against EFL writers within this dataset." Originality classified every news text correctly but missed four student essays, a difference the authors describe as a borderline trend (one-sided p = 0.058), in the same direction as earlier research on detector bias against non-native writers.

Read that together with the size of the test. Four essays out of 48 is roughly one in twelve. That is not zero, and a result above 20% is exactly the kind a student might be asked about. It is also a comparison with news journalism, not with academic essays by native speakers.

Mixed human and AI text was the hardest case

On the half-human, half-AI texts, both detectors mostly failed to put them in the middle band. Turnitin placed 31% of them there and Originality 2%. The authors' reading is that the detectors "appear calibrated primarily for binary distinctions (AI vs. Human) and struggle to identify partial or mixed AI involvement."

Overall accuracy across all four groups was 0.61 for Turnitin and 0.69 for Originality.

Science writing was harder to classify

Both detectors were much less accurate on science texts than on humanities texts:

AccuracyScienceHumanities
Turnitin0.510.86
Originality0.580.96

These figures cover all four groups together, so they measure errors in every direction, not only human writing scored as AI. The authors suggest that features common in scientific writing, "such as lexical density, technical terminology, and formulaic sentence structure, may be more difficult for detectors to distinguish from AI-generated text."

A different 2026 study, of Pangram and GPTZero on published abstracts, found recent political science and theology abstracts flagged more often than chemistry and computer science. That study measured something else, on different detectors, which we cover in does AI polishing get your writing flagged. The two results do not add up to a rule about which subjects are safer.

Length mattered too, though not in a straight line. Across all four groups, both detectors were most accurate on the shortest texts (300 to 330 words) and least accurate on the middle band (450 to 550 words): Turnitin 0.87, 0.56 and 0.68 for short, medium and long texts, Originality 0.96, 0.63 and 0.84.

What this study does not tell you

  • Two detectors, tested in early 2025. Only Turnitin and Originality were tested, between January and May 2025, and detectors change with updates.
  • One group of students. The EFL essays came from one foundation programme at one university.
  • News, not academic prose, as the comparison. The professional texts were BBC articles.
  • Fixed mixtures. Every hybrid text was about 50/50, which the authors note real student writing rarely is.

Their conclusion is that "neither detector achieves the level of reliability required for high-stakes decision-making."

If your own writing was flagged

  1. Find out which band you are in. In this study, a score of 20% or less counted as human. Turnitin shows an asterisk instead of a number for low scores; see what the asterisk on a Turnitin AI score means.
  2. Keep your drafts and version history. They show how the text was written, which a score cannot.
  3. Look at which passages were highlighted. Those are the ones to be ready to explain.
  4. Raise the evidence. Studies like this one, and the earlier research on non-native writers in why non-native English writers get flagged more often, support the point that a detector score alone is not proof.

Where HumanPen fits

If you have a Turnitin or iThenticate report and the flagged passages need rewriting, that is what HumanPen's humanize option does. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file. Check your course's rules on AI assistance first, and read what comes back before you submit it.

Frequently asked questions

Is Turnitin biased against non-native English writers? In this dataset, Turnitin scored EFL student essays and BBC news texts about the same, with no significant difference. That is one test of one group of students, not a general answer.

What did "flagged" mean in this study? A score above 20%. Scores from 21% to 79% counted as Hybrid and 80% or more as AI.

Were the student essays really written without AI? They were submitted before August 2022. The authors say tools like ChatGPT were not yet available to those students, and text generators that existed then required restricted access and technical setup.

Which detector did better? Originality had higher overall accuracy (0.69 against 0.61). On student essays the two were level, each scoring 4 of 48 above 20%, while Originality got every news text right. Neither handled mixed texts well.

Sources

KEEP READING