Does Turnitin detect Claude, Copilot or Gemini?
Switching from one assistant to another assumes the detector recognises whose model produced the text. Turnitin's FAQ, update pages and report documentation describe coverage, releases and output at different levels. Keeping those levels apart changes what can honestly be concluded.
HumanPen Team
· 22 min read
The short answer
Claude and Gemini versions appear in a list of models Turnitin says its English detector "can detect content from". Copilot is not named in the five pages we reviewed, but that is neither an exemption nor evidence that a particular Copilot submission will be detected. Product Updates and general release notes also name models in language-specific updates, so the FAQ is not Turnitin's only public source of model names. None of the reviewed material provides a stable per-model odds table. The report documentation instead describes probabilities calculated over overlapping passages of text, with output classified as likely human or likely AI-generated.
The sources are five named pages read on 18 August 2026: the AI Writing Report guide, the detection FAQ, the AI writing detection model change log, the general Turnitin release notes, and Product Updates. Conclusions below stay inside those pages and the cited sections. They do not claim to exhaust Turnitin's website, and the article does not publish character totals or word-hit counts from pages that update in place.
One thing this is not arguing: that switching models does nothing. The reviewed sources do not provide a number that would let us predict that comparison, and neither do we.
The FAQ's coverage list
The list sits in Turnitin's AI writing detection FAQ, and it appears on that page twice: once inside the answer to "How does it work?", and again under the question "Which AI writing models can Turnitin's technology detect?". Both copies open identically.
"Currently, Turnitin's AI writing detection model for English submissions can detect content from GPT-4o (released 2024-05) ..."
The verb is worth a second look. "Can detect content from" is a statement about coverage. It is not a note about what once went through a test, and a lot of writing on this subject quietly softens it into the second thing. On 18 August 2026 the sentence named multiple GPT, Gemini and Claude versions, alongside entries from LLaMA, Mistral, Deepseek, Nova and Grok plus o1-mini. That is a dated snapshot of the families present, not a permanent total.
Each entry carries a release date in brackets. Ignore those. The two copies of the same list on the same page disagree with each other on at least two of the dates, and they still disagreed when we checked on 18 August 2026, so whichever one you read, the page itself is telling you the dates are not maintained.
The sentence then closes with "and tools based on these LLMs as well", and the sentence after it says Turnitin will "continue to expand our detection capabilities to other models in the future".
What Copilot's absence does and does not show
People usually search with a product name rather than a model name. In the five pages named above, we did not find a Copilot-specific coverage statement, accuracy claim or result prediction. That is a bounded finding about those pages, not a claim about every Turnitin page.
The FAQ closes its model list with "and tools based on these LLMs as well". That general clause may be relevant when a tool uses a listed model, but this article has not established which model a particular Copilot product, mode and date used. Without that mapping, the clause cannot tell us whether one Copilot submission will be flagged. The missing product name proves neither exemption nor detection.
A coverage list is not a per-model ranking
The English FAQ names models in a coverage statement, but that answer does not attach a detection rate, accuracy figure or difficulty ranking to each entry. The one figure it does attach to detection is an overall false-positive aim: "under 1% for documents with over 20% of AI writing". The figure applies to the detector overall; it does not rank Claude against Gemini or any other pair.
Other Turnitin pages do publish model names. Product Updates named GPT and Gemini versions in its 5 May 2026 Spanish-model update. The general Turnitin release notes added an Arabic-model update on 18 August 2026 that names versions from several model families. The separate AI writing detection model change log summarizes some releases with less detail. Those pages directly disprove the idea that the FAQ is the only public place where Turnitin names models.
What those named updates still do not provide is a stable table of per-model odds for an arbitrary submission. A release coverage note and a prediction for one file are different claims.
Where the number actually comes from
The FAQ describes the pipeline in one paragraph, and it is short enough to read in full:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score."
Two categories in that last clause. Likely human, likely AI-generated. Whatever anyone tells you about a detector recognising a particular model's fingerprint, this is the vendor's own account of what its classifier emits, and there is no third slot in it. Turnitin also says its model is not built out of named metrics: it "is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions", and "instead, it learns statistical patterns from our training data". What that distinction changes in practice is worked through in what AI detectors measure.
If you were picturing a bank of detectors, one trained per generator, the FAQ describes the opposite direction of travel in the sentence immediately after the model list:
"In July 2026, we updated our model architecture to consolidate a multi-model ensemble into a single model. This update improves and simplifies the AI writing report, maintaining a less than 1% false positive rate."
The FAQ describes one classifier producing text-level probabilities. Separately, it publishes a model coverage statement. The July architecture sentence appears in the FAQ but not in the AI writing detection model change log. That mismatch shows why the page type has to be named instead of treating every update page as one exhaustive release history.
The vendor removed the one distinction people were reading as attribution
On 4 August 2026 Turnitin merged two report categories into one. Purple highlighting, which had marked AI-generated text that was also AI-paraphrased, stopped being shown. The reasoning it gave on its product updates page is the most direct thing the company has published about tool attribution, and it is describing its own users:
"Reduce the risk of unwarranted academic misconduct reviews by removing distinctions that could be read to imply more certainty about the specific tool or workflow used to produce a submission"
A distinction the vendor added in 2024 came back out in 2026 because readers were pulling tool identity out of a report that could not carry it. Note what did not change. The release note announcing the merge says the model "will continue to detect likely AI generated content that may have been further modified by AI paraphrasers or bypassers". Fewer colours in the interface, same detection claim underneath.
There is a practical edge to this. If you are looking at a screenshot with purple in it, that report was not generated under the current rules, because Turnitin says the update applies to newly submitted files and that existing submissions must be resubmitted to get the updated report. Date a report before reasoning from it. How to read a Turnitin AI writing report covers what the current figure includes.
What the same pages do say moves the number
Length, for one, and in a direction that surprises people. Short pieces get blunter results rather than safer ones:
"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap. This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."
That is a property of the document, not of whoever drafted it. The same holds for the properties Turnitin names when describing its own errors:
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."
Structural variation, repetition, ideas that do not develop. Every item is something you could measure with the source hidden. None of them is a vendor. That does not add up to "switching models is pointless", and it would be dishonest to dress it up as that, since the pages simply do not address the comparison. It does mean that the variable people spend the most time on is the one their own documentation says least about. If you have been flagged for something you wrote yourself, preparing a response is a better use of an afternoon.
What a company in our position can honestly say
We sell a document rewriting tool, so read this section with that in mind. We do not promise a detection outcome, and the reason is not modesty. Detection systems update. The public Product Updates page records the 4 August report change, which is enough to show that the report itself can change without making any claim about who received a separate notification.
Scope and cost can be described honestly, so those are what HumanPen describes. A Turnitin or iThenticate report is really a list of coordinates, and handing it over with the file confines the rewriting to those coordinates; sentences outside them keep the wording you gave them. Credits count how much prose was changed, not how large the file was. Eligible results can continue lowering AI for free.
There is also a real case for not using anything of this kind, and it is set out separately in when a document-level rewriting tool is the wrong tool. Six situations, each quoted from our own pages, where the honest answer is your own hands.
Frequently asked questions
Copilot is not named. Does that help? The five pages reviewed do not give a Copilot-specific conclusion. The FAQ's general clause about "tools based on these LLMs" may apply when a tool uses a listed model, but we did not establish the model behind a particular Copilot product, mode and date. A missing name cannot predict one submission's result.
Is one of the listed models safer than the others? The reviewed pages do not provide a stable per-model ranking. Product Updates and release notes sometimes name models when describing coverage changes, but that is not a probability table for the next file. A percentage from someone's own submissions is a different kind of evidence and should be read as one. The general version of this is why the same text scores differently on every detector.
Does any of this apply to a Spanish or Japanese submission? Those use separate detection models, and the published detail differs by page. Product Updates names GPT and Gemini versions in its 5 May 2026 Spanish update, while the separate AI writing detection model change log summarizes that release without the same list. The general Turnitin release notes are a third source, not another name for either page. More on the language boundaries is in can Turnitin detect AI in non-English submissions.
Can a marker see which model I used? The current report documentation we reviewed describes likely-human and likely-AI output, not a model-name field. Since 4 August 2026 the report has shown one AI-generated category, and the company's stated reason for the merge was to remove distinctions "that could be read to imply more certainty about the specific tool or workflow used to produce a submission".
I used an older version of a model that is on the list. Different result? The reviewed pages do not let us predict that. Versions appear separately in the FAQ list, each with a bracketed date, but the two copies contradict each other on some dates and do not attach a result figure to each version. A dated coverage update still would not predict one file.
KEEP READING