Machine-Translated Your Own Writing? How Often Pangram Called It AI in a 2026 Test

You drafted in Chinese, Arabic or Spanish and let a translator carry the text into English. Then the worry starts: a machine produced that English, so won't a detector read it as machine-made? A preprint posted on 5 October 2026 measured this for one detector. The number is small. The conditions around it matter more than the number, and none of them answer a separate question, which is whether your course lets you translate in the first place.

HumanPen Team

· 11 min read

How often does an AI detector flag writing you translated yourself?

In one 2026 test, rarely. Researchers took human-written web pages in 24 languages, all from before 2022, and had Gemini 3.7 Flash translate them into English as faithfully as it could. Pangram 4 called 26 of the 1,197 translations AI, or 2.17%. It had called none of the 1,198 originals AI. When a model had done the writing, though, the result flipped: pages that Gemini had rewritten from the person's own outline were flagged 1,161 times out of 1,197 (97.0%) once they were in English.

This was Pangram only. Turnitin, GPTZero and other detectors weren't tested, so the 2.17% doesn't carry over to them. The paper is a preprint and hasn't been peer reviewed, the texts were web pages rather than essays, and the translation was deliberately literal. Whether a detector flags your translation is also a different question from whether you were allowed to translate. Several universities, and JCQ, the body that sets coursework rules for GCSEs and A-levels, have written rules on exactly that.

The study is "IdeaLens: Detecting AI Ideas in Long-form Writing" by Rishanth Rajendhran and seven co-authors at the University of Maryland, Google DeepMind and Simon Fraser University, posted to arXiv on 5 October 2026. The translation test is a side experiment in its appendix. The paper's main subject is a research detector that tries to tell whether a document's ideas came from a person or a model. The authors thank Pangram, the detector measured here, for a research credit award.

What the researchers translated

The human texts came from FineWeb2, a large public collection of web pages. The authors kept pages written before 2022, which makes it very unlikely that a chatbot wrote them, and only in eight long-form types: academic writing, creative writing, knowledge articles, news articles, nonfiction, personal "about" pages, personal blogs and user reviews. For the translation test they took 50 pages in each of 24 languages. The list runs from Chinese, Spanish, German, Japanese, French and Russian through Arabic, Hindi, Persian, Vietnamese and Ukrainian to Amharic, Yoruba, Sinhala and Tamil.

Each human page came with two AI versions in the same language, both written by Gemini 3.7 Flash:

  • Topic only. The model got a one-line description of the page type and subject, so everything in the text was its own.
  • From the person's outline. The model got the page's full outline and was told to write out every item in order and add nothing. The person supplied every point. The model wrote every sentence.

All three versions were then translated into English, again by Gemini 3.7 Flash. Its instructions asked for "nothing added, nothing removed, nothing improved" and told it to "Keep the register, repetition, fragments and clumsy phrasing." That's closer to a careful human translator than to pasting a draft into a chatbot and asking for better English. Pangram 4 scored every text before and after translation.

What Pangram said, before and after translation

TextFlagged as AI in the original languageFlagged as AI after translation into English
Written by a person0 of 1,198 (0%)26 of 1,197 (2.17%)
Model wrote it from the person's full outline616 of 1,200 (51.3%)1,161 of 1,197 (97.0%)
Model wrote it from a one-line topic942 of 1,200 (78.5%)1,200 of 1,200 (100%)

"Flagged" has a specific meaning here. Pangram 4 splits a text into shares it judges AI, AI-assisted and human. The authors counted a text as AI when the AI share plus half the AI-assisted share came to at least half. A translated page where Pangram marked only a small part as AI wouldn't be in the 26, and the paper doesn't say how many pages like that there were.

So about one translated human page in 46 got an AI verdict. Two other rows are easy to misread.

The 97% is not a false alarm about human writing. The paper files those flags under false positives, because the paper is about ideas and in that row the ideas were the person's. But a model wrote every sentence, and wording is what Pangram is built to read. The dataset's own description says so directly: "An idea detector should not flag level 5; a prose detector should." (Level 5 is the from-your-outline version.) If you hand a chatbot your notes or outline in your own language and let it write, the English you get back is the chatbot's prose, and in this test Pangram caught nearly all of it.

In this test, translation didn't hide AI text. It exposed more of it. In the original languages Pangram missed a fair share, 258 of the 1,200 topic-only texts in this test, and more in languages with less text online: across the paper's full multilingual set, it caught 58.0% of topic-only texts in the ten languages the authors group as low-resource. After translation it flagged all 1,200 topic-only texts. Writing with a chatbot in your own language and then translating got caught more in this test, not less.

A second check in the same paper points the same way. CoCoNUTS, a dataset from another research group, contains peer reviews written in English in 2022 or earlier. Some were run through Chinese and back into English by language models, and Pangram 4 flagged 0.1% of those 1,493. Others from the same collection were polished by language models, and Pangram flagged 61.2% of those 1,497. A round trip through Chinese isn't the same thing as your draft. Still, across both datasets the AI verdicts piled up where a model wrote or improved the sentences, not where it translated them faithfully.

What this doesn't tell you

  • Anything about Turnitin. No Turnitin figure comes out of this study, and a Pangram verdict and a Turnitin percentage aren't the same kind of measurement (more in Pangram vs Turnitin: Why the Two AI Scores Don't Line Up). A common argument goes that translators and chatbots run on the same kind of language model, so translated English should look machine-made to any detector. For Pangram and faithful translation, this measurement mostly didn't bear that out. This study didn't test Turnitin, so it says nothing about how Turnitin treats translated text. An older test from 2023 did include Turnitin among 14 detectors: across all of them, overall accuracy was 96% on human texts written in English and, in the authors' words, "dropped by 20%" on nine human texts machine-translated into English with DeepL or Google Translate. That is an average over all 14 tools, not a Turnitin number, and detectors have changed since. Does Translating Chinese to English Get Flagged as AI by Turnitin? goes through the patterns Turnitin's own documentation links to false positives.
  • Your language. The translation results are published only as a total across 24 languages, 50 pages each. There's no Chinese-only or Arabic-only rate for this test.
  • Your translator. One model with a strict fidelity prompt. DeepL, Google Translate and a chatbot told to "make it sound academic" weren't tested. The polished-review result suggests that last one is a different situation.
  • Essays. These were blogs, news stories, reviews and similar pages, not assignments written to a rubric.
  • Peer review. It's a preprint, not yet peer reviewed, and Pangram supplied research credits. The paper doesn't say exactly when the texts were scored, and detectors get updated.

A low flag rate is not permission

The study can't answer this part. Universities write rules about translation tools specifically, and those rules apply whether or not anything gets flagged. Three current university pages, and JCQ's coursework rules:

  • University of York. Its assistance policy lists, as unacceptable for summative work, "Using translation tools or services to translate text in whole or significant sections (oral or written), which is then submitted as summative work". Its student guidance accepts translation "As a dictionary" at the level of a word, collocation or phrase, and warns that "translated text which is at sentence level or beyond risks submitting work of false authorship."
  • University of Huddersfield. "If you translate large portions of text written initially in your native language and submit this for assessment, this may be considered to be misconduct, as it raises the question of authorship of the translated version." The same regulation asks you to reference translation tools and says a tutor who questions your language may ask you to summarise the content or explain the grammar and vocabulary you used. The next clause, 10.2.3, adds: "We do not expect you to use AI tools to contribute to the completion of your assessment or online exam unless you have been explicitly told that you can." That may matter if your translator is a chatbot.
  • University of Queensland. Its library guide says to check your course profile before using machine translation, rules it out for in-person assessment, and says: "You must acknowledge that you have used machine translation in your assessment. Failure to acknowledge the use of machine translation can result in Academic Misconduct."
  • JCQ (GCSE, A-level and EPQ coursework). The 2026-27 JCQ instructions for GCSE and A-level non-exam assessment and for the EPQ now say: "AI translation tools must not be used to generate work for assessment."

Set those next to the study. At York, an essay translated whole from your own draft is on the unacceptable list even if Pangram would pass every page of it. At Queensland the same essay can be fine, as long as your course allows it and you say you did it. The detector's verdict changes neither outcome.

If you draft in your own language

  1. Find the rule for this assignment, not just the university's general page. York lets departments switch translation off for particular assessments, and Queensland points you to the course profile. If the brief says nothing, ask in writing: "I write my first draft in [language] and translate it with [tool]. Is that allowed for this assignment, and how should I acknowledge it?" Keep the reply.
  2. Acknowledge it in the form your school asks for. Queensland's guide gives sample wording you can adapt: "This work was post-edited by the author, after being translated from <insert language> to English using <insert tool>."
  3. Keep your first-language draft, with dates. It's the clearest record that the ideas and structure were yours before any tool touched them. A file in OneDrive or Google Docs keeps version history. A Word file sitting on your desktop doesn't. How to keep version history in Word, Google Docs and Overleaf walks through all three.
  4. Keep the model out of the writing step. In this study, AI verdicts stayed rare when a model only translated, and climbed when a model wrote from the person's outline or polished their sentences. York's policy separately lists "Generating or re-writing (including shortening or summarising) any of the student's sentences or sections of work" as unacceptable. "Translate this and make it sound more academic" is really two requests, and the second is the one both the detector data and that rule are about. Some courses allow more; check yours.
  5. Make sure you could explain every sentence. Read the English against your draft until you could say what each sentence means, and why it uses the words it does. Huddersfield's regulation describes exactly that kind of check.

Frequently asked questions

Does this mean Turnitin won't flag my translated essay? No. Turnitin wasn't part of the study. Apart from the authors' own research models, the only detector run on the translated pages was Pangram 4.

Was Chinese one of the languages? Yes, Chinese was one of the 24. But the translation results are reported only as a combined total, so there's no separate rate for Chinese.

Is writing straight in English safer than translating? This paper can't settle that. In a separate set of 3,909 essays by English learners, Pangram 4 flagged 0.0% (rounded to one decimal place). The translated web pages came in at 2.17%. Those are different kinds of text, so the two numbers aren't a head-to-head comparison. Both are low.

What is IdeaLens, and should I worry about it? It's the authors' research detector, built to judge where a text's ideas came from rather than who wrote the words. On the translated human pages it flagged 3 of 1,197. The authors have released it publicly. The paper doesn't say any school uses it.

My translated essay has already been flagged. What now? Ask which detector was used and for the report itself, not just a number. Bring your first-language draft, your translation history and any version history. Then check the assignment rules, because whether translation was allowed and whether you acknowledged it will come up whatever the detector said.

Sources

KEEP READING