Does Professional English Editing Make Your Paper Look AI-Written? A 2026 Study

If you write in English as a second language, you may be wondering whether sending your paper to a native-speaker editor will make an AI detector suspicious. A 2026 study tested exactly that on more than 135,000 real manuscripts, before and after human editing. The answer depended mostly on the detector. Here is what it found, who ran it, and why it cannot tell you what Turnitin would do.

HumanPen Team

· 6 min read

Does professional editing make a paper look AI-written?

Not consistently, in this study. Across 13 open-source AI detectors, human editing raised the AI score on some and lowered it on others, and the change reached practical significance on only 7 of them. The detector made far more difference than the editing: on the same unedited manuscripts, false-positive rates ran from 0.0% to 100%. No commercial detector, including Turnitin, was tested.

The study is "Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing", by Hyeonchu Park and Bugeun Kim of Chung-Ang University and Gahye Jeong of Wordvice, posted to arXiv on 27 August 2026 and listed there as presented at EMNLP 2026. The documents came from Wordvice, a professional English editing service. One author, Gahye Jeong, is Wordvice's CTO according to the company's own blog, which also describes Wordvice as "the developer of the Wordvice AI Detector".

What the researchers compared

The data was 135,389 pairs of documents from 2018 to 2025, spanning more than 40 academic disciplines. Each pair was a document written by a non-native English speaker and the same document after a native-speaker editor had worked on it. Just over half (52.9%) came from academic editing and 37.6% from admissions editing; application essays were the largest single subject area, with 22,639 documents. The rest were business editing, TOEFL writing, translation and other work. Because the author, topic and content stayed the same, any change in a detector's score came from the editing.

The editing itself was done by people. The service's guidelines, as the paper describes them, allow AI grammar checkers only for mechanical errors, and "AI-generated rewriting, paraphrasing, or stylistic generation is explicitly prohibited." The edits were real but modest: in a 6,000-pair sample, the median document changed length by only 6.5 words, and average sentence length fell slightly, from 24.35 to 23.51 words.

The 13 detectors were all publicly available research tools, run in their standard settings: statistical methods such as GLTR, zero-shot methods such as Fast-DetectGPT and Binoculars, and trained classifiers such as RoBERTa and RADAR.

Which detector mattered more than anything else

On the manuscripts written before 2023, which the authors treat as human-written, the share wrongly classified as AI depended almost entirely on the detector:

Detector typeExamplesUnedited manuscripts classified as AI
Token statisticsLog-rank, Entropy, GLTR99.8% to 100%
Trained classifiersRoBERTa, RADAR93.1%, 88.3%
Trained classifierMAGE16.7%
Zero-shotFast-DetectGPT, LastDE+25.2%, 19.9%
Zero-shotDetectLLM-LRR, DiVeye, BiScope, Binoculars0.0% to 0.6%

The same human-written documents were classified as almost entirely AI by some tools and almost entirely human by others. The authors conclude: "These findings suggest that AI detectors do not represent a unified measure of AI-generated content; rather, they rely on a mix of generation-origin and linguistic-quality signals." They also caution that these absolute rates depend on the thresholds they set.

What editing changed

After editing, the direction of the change depended on the detector. These figures cover the full 2018–2025 data, not only the pre-2023 documents in the table above:

  • Fewer false positives after editing: MAGE (down 10.7 percentage points), RADAR (down 4.7), BiScope (down 0.5).
  • More false positives after editing: LastDE+ (up 5.8), Fast-DetectGPT (up 4.6), RoBERTa (up 1.7).
  • Unclear: DiVeye, whose figures point in different directions in different tables of the paper.
  • Little or no change: the remaining detectors, within about 0.1 points.

The authors regard these before-and-after changes as sturdier than the absolute rates: "the relative patterns across editing conditions provide more robust evidence of detector sensitivity to linguistic refinement."

The more heavily a manuscript was edited, the further its score moved, in whichever direction that detector already leaned. The effect was also weaker in 2023 to 2025 than in 2018 to 2022.

What this study does not tell you

  • No commercial detectors. In the authors' words, "the analysis does not cover future systems or proprietary commercial detectors." Turnitin, GPTZero, Originality and Pangram were not tested, and the paper does not mention Turnitin.
  • Mostly academic editing, but not only. The authors describe the data as predominantly academic manuscripts, yet more than a third was admissions editing, and the paper does not report results separately by type. They add: "Detector behavior may differ for student essays, journalistic articles, creative writing, social media content, and other genres with distinct linguistic and stylistic characteristics."
  • Human editing only. The editing in this data was done by people under rules that prohibit AI rewriting. It says nothing about AI editing tools.
  • No AI-written controls. In the authors' words, "the study does not include AI-generated control texts subjected to comparable editing procedures."
  • Who ran it. The data and one author came from Wordvice, which sells editing services and, by its own description, develops an AI detector.

What to take from it

  1. Weigh editing on its merits, not on these detectors. For the 13 research detectors tested here, the study gives no reason to drop a human editor out of fear of AI flags: editing moved most scores only slightly, and in both directions, though three detectors did flag more after editing. Whether the same holds for Turnitin was not tested. Check whether your journal or course asks you to declare editing help.
  2. Keep both versions. Your original draft and the editor's tracked-changes file show exactly what a person changed, which a detector score cannot.
  3. Ask which detector produced a score. Detectors in this study treated the same human writing very differently, though the authors note the exact rates depend on their thresholds. For earlier research on non-native writers, see why non-native English writers get flagged more often; for a 2026 test of Turnitin itself on EFL student essays, see does Turnitin flag non-native English writing.

Where HumanPen fits

If you have a Turnitin or iThenticate report and the flagged passages need rewriting, that is what HumanPen's humanize option does. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file. Check your journal's or course's rules on AI assistance first, and read what comes back before you submit it.

Frequently asked questions

Will Turnitin flag my paper because a native speaker edited it? This study cannot say; it did not test Turnitin or any other commercial detector.

Did editing make papers look more human or more AI? Both, depending on the detector. Three gave fewer false positives after editing and three gave more; DiVeye's figures are inconsistent across the paper's tables, and the rest barely moved.

Does this apply to AI editing tools? No. The editing in this data was done by people, under guidelines that prohibit AI rewriting. For a 2026 study of AI polishing, see does AI polishing get your writing flagged.

Why did some detectors flag nearly every manuscript? The authors found that detectors based on token statistics had the highest false-positive rates on these human-written texts, while likelihood-ratio zero-shot detectors stayed close to zero. They also note that absolute rates depend on the thresholds used.

Sources

KEEP READING