Does AI Polishing Get Your Writing Flagged? What a 2026 Study Measured
You wrote it yourself and used AI only to smooth the language, and now a detector says AI. A 2025 study had already shown that even very light AI polishing raises detection rates. In August 2026 a University of Notre Dame team measured it again on real published abstracts, field by field. Here is what they tested, the numbers, the limits the authors put on them, and what to do with a flag of your own.
HumanPen Team
· 7 min read
Does AI polishing get your writing flagged?
Often, on the two commercial detectors this study tested. Researchers asked an AI model to rewrite 642 published abstracts "to sound better and more human-like". Pangram flagged 64 to 80% of the polished versions and GPTZero 38 to 49%. The same detectors flagged none of the original abstracts written in 2013 to 2015. Turnitin was not tested. The authors' conclusion is that "detector scores should not serve as standalone misconduct evidence."
The study is "Why AI Detection Fails for Academic Integrity", by four researchers at the University of Notre Dame, posted to arXiv on 6 August 2026 and written for the ACM AI Leadership Summit (AILS '26).
It is not the first to look at this. A 2025 study by Saha and Feizi at the University of Maryland, published in the Findings of ACL 2025, found that after "only extremely minor edits" by one AI model, 42.56% of human-written samples were detected as AI by Pangram and 64.71% by GPTZero, using a threshold set for each detector. The two studies used different models, thresholds and texts, so their numbers sit side by side rather than add up. They point the same way.
What the researchers did
They took 642 English abstracts of published papers from four fields: chemistry, computer science, political science and theology. 306 came from 2013 to 2015, before tools like ChatGPT existed, and 336 from 2023 to 2025.
Each abstract was rewritten by Gemini 3 Flash. The version that matters here is the one the authors call a "refine", which they treat as "a proxy for guideline-compliant AI assistance". The instruction the model received began:
"You are an expert scientist who excels in writing papers. Your task is to rewrite a given paper abstract to sound better and more human-like."
The authors are clear about what kind of edit that produces: "Our refine prompt is closer to light generative rewriting than to grammar-only editing; assisted-writing flag rates should be read accordingly." Keep that in mind before comparing it with what you did.
Every original and every rewrite was then scored by Pangram 3.2 and GPTZero. According to the authors' published code, the score used for GPTZero is the probability in its "AI" class, and anything at 0.50 or above counted as flagged.
The numbers
| What was scored | Pangram flagged | GPTZero flagged |
|---|---|---|
| Original abstracts, 2013–2015 | 0.0% | 0.0% |
| Original abstracts, 2023–2025 | 14.9% | 8.9% |
| AI-polished versions of the 2013–2015 abstracts | 64.4% | 37.6% |
| AI-polished versions of the 2023–2025 abstracts | 80.1% | 48.5% |
Two cautions come from the authors themselves. The 2023–2025 originals may already have had some AI help that nobody recorded, so those flags are "not confirmed false positives". And the polished versions are a different thing from human text: the paper says their flag rate "must not be abbreviated as FPR", so we do not call it a false-positive rate here either. What it measures is how often a human-written abstract is flagged once an AI has polished it.
Moving the threshold did not change the picture much. Between 0.4 and 0.6, Pangram flagged 80 to 85% of the polished 2023–2025 versions, and GPTZero 48 to 49%.
How much the field mattered
On the 2023–2025 abstracts, the rates depended heavily on the subject:
| Field | Originals flagged (Pangram / GPTZero) | Polished versions flagged (Pangram / GPTZero) |
|---|---|---|
| Chemistry | 2.1% / 1.1% | 73.7% / 46.3% |
| Computer science | 9.6% / 6.0% | 69.9% / 36.1% |
| Political science | 24.1% / 15.2% | 86.1% / 54.4% |
| Theology | 26.6% / 15.2% | 92.4% / 58.2% |
On the unedited originals, the gap between the STEM fields (chemistry and computer science) and the others (political science and theology) was statistically significant. The authors also measured which features of the text went with higher scores. The two strongest were the share of long words and the share of words from the Academic Word List, a standard list of vocabulary common in academic writing, with correlations of about 0.30 to 0.35. A higher share of numbers and other non-alphabetic tokens went with lower scores. The authors state that these "associations are correlational and do not establish causal mechanisms."
These are measurements of Pangram and GPTZero, not Turnitin. In the three Turnitin help pages on AI detection that we searched in why AI detectors flag well-written essays, there was no language about grading writing quality. The two findings are about different tools and do not contradict each other.
What the study does not tell you
The authors list the limits themselves, and they matter for applying this to your own work:
- Published abstracts, not student essays. In their words: "We analyze English published abstracts, not student essays; detector behavior on classroom genres may differ."
- Two commercial detectors, not Turnitin. The commercial detectors were Pangram 3.2 and GPTZero, alongside a supplementary AI-model baseline. The paper does not mention Turnitin. Pangram has released newer versions since, so a Pangram result today may not come from the same model.
- One model, one kind of polish. Every rewrite came from Gemini 3 Flash with the instruction above. A spelling-and-grammar pass was not tested separately.
- Four fields, English only. A low rate for chemistry here does not mean chemistry students are safe with other detectors or other kinds of writing.
- Pangram provided API access for the study, according to the acknowledgments.
The paper also put the AI rewrites through a commercial humanizer, after which fewer than 4% were still flagged by either detector. Set beside the 38 to 80% flag rates for honest polishing, that is the other half of the authors' argument that detector scores cannot stand on their own. Their conclusion: "Integrity programs therefore need transparent AI-use norms and process evidence, not standalone detector scores."
If your own writing was flagged after polishing
- Find out which tool produced the score. If it was Turnitin, these numbers are not about it. Read your instructor's report for what it actually highlights.
- Check what your rules allow. Many publisher policies treat language editing differently from generating content, and some want it declared; our guide to publisher AI disclosure policies compares them. For coursework, the assessment brief and your university's policy decide.
- Keep the version from before the polish. The paper's practical advice is that "Any flag must be paired with drafting history". A draft in your own words, with its version history, is that history.
- Look at which passages were flagged, not only the percentage. Those are the passages you should be ready to explain: how you wrote them, and what the AI changed.
Where HumanPen fits
If you have a Turnitin or iThenticate report and the flagged passages need rewriting, that is what HumanPen's humanize option does. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file. Check it against your course's rules on AI assistance first, and read what comes back before you submit it.
Frequently asked questions
Does using Grammarly count as AI polishing in this study? Not directly. The study's polish was a rewrite by Gemini 3 Flash, which the authors describe as "closer to light generative rewriting than to grammar-only editing". A spelling-and-grammar pass was not tested. Turnitin has said separately that it does not target Grammarly's spelling, grammar and punctuation fixes; see does Turnitin detect Grammarly.
Can I use this study if I am accused of using AI? It supports a general point, that these detectors flag a large share of AI-polished human writing and that the authors say scores "should not serve as standalone misconduct evidence". It does not describe Turnitin or student essays, and it cannot show what happened with your text. Your drafts and version history do that.
Why would unedited 2013–2015 abstracts score 0% but 2023–2025 ones score higher? The paper cannot say which cause it is. In its own words, the 2023–2025 window "cannot separate model vintage from unobserved real-world AI use on original abstracts". That is why it reports those figures as flag rates, not false positives.
Does GPTZero label polished text differently? GPTZero has an "AI Polished" label under its Mixed result; we went through what it means in GPTZero's Mixed and AI-Polished labels. According to the authors' code, this study used the probability in GPTZero's "AI" class, so a text GPTZero put in Mixed, including AI Polished, counted as flagged only if its AI-class probability still reached 0.50.
Sources
- Jonathan A. Karr Jr., Grigorii Khvatskii, Ting Hua and Nitesh V. Chawla, Why AI Detection Fails for Academic Integrity, arXiv:2608.11256 (v1 6 August 2026, v2 20 August 2026), AILS '26; read 27 September 2026 (Tables 1–6, Appendices A, B and D), with the authors' published code.
- Shoumik Saha and Soheil Feizi, Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing, Findings of ACL 2025 (arXiv:2502.15666); read 27 September 2026.
KEEP READING