Does Turnitin Detect GPT-5 and Other New AI Models?
Students wonder whether the latest AI models can bypass Turnitin. The FAQ publishes a list of models it can detect content from, and says it will continue to expand detection to other models. The FAQ also says detection capabilities will change over time. Here is what the documentation actually says.
HumanPen Team
· 12 min read
The Short Answer
Turnitin's FAQ publishes a list of AI writing models it can detect content from. The list includes GPT models, Gemini, Claude, LLaMA, Mistral, Deepseek, Nova, Grok, and o1-mini, among others. The official verb is "can detect content from." The list ends with: "and tools based on these LLMs as well." The FAQ also states: "We will continue to expand our detection capabilities to other models in the future." The FAQ acknowledges that detection capabilities are not static: "As we iterate and develop our model further to better detect newer LLMs, it is likely that our detection capabilities will also change, affecting the AI percentage." Whether a specific model like GPT-5 appears on the list depends on when the FAQ was last updated. The FAQ does not promise real-time coverage of every new model release, but it does promise to keep expanding.
What the Model List Says
The FAQ provides a list of models it can detect content from. The introductory sentence uses the verb "can detect content from," followed by a comma-separated list of models. The list includes GPT, Gemini, Claude variants, LLaMA, Mistral, Deepseek, Nova, Grok, and o1-mini. The list concludes:
"and tools based on these LLMs as well."
The next sentence: "We will continue to expand our detection capabilities to other models in the future."
This means the list is not exhaustive. Tools built on top of listed models are also claimed to be detectable. The promise to expand means new models may be added as the detection model is updated. Does Turnitin detect Claude, Copilot or Gemini works through the named entries on that list.
Detection Capabilities Change Over Time
The FAQ explicitly acknowledges that detection is not a fixed target:
"As we iterate and develop our model further to better detect newer LLMs, it is likely that our detection capabilities will also change, affecting the AI percentage."
The next sentence: "However, for a submitted document, the AI percentage will change only if it's re-submitted again to be processed."
This has two implications. First, the detection model is being updated to target newer LLMs. Second, already-graded documents do not automatically get re-evaluated when the model changes. If your paper was scored before a model update, the score reflects the detection capabilities at that time, not the current ones — which is why still flagged after a Turnitin recheck argues for comparing reports rather than scores.
How the Detector Actually Works
The detection model processes text by splitting it into overlapping segments and assigning probability scores:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."
The model is not built on surface-level metrics like burstiness or perplexity. The FAQ states: "Our model is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions." The next sentence: "Instead, it learns statistical patterns from our training data."
The same FAQ also says: "Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers." So while the model does not calculate burstiness or perplexity as named metrics, it does classify text based on word probability patterns learned from training data.
What the Detector Might Miss
The FAQ acknowledges a trade-off between false positives and false negatives:
"In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document."
The FAQ elaborates: "For example, if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."
This means the detector is tuned to avoid false positives even at the cost of missing some AI text. A new model that produces text closer to human word probability patterns would be harder to detect, at least until the detection model is updated to cover it. The FAQ's promise to "continue to expand our detection capabilities" is the counterweight, but there is an inherent lag between a new model's release and detection coverage. The same lag question comes up for rewriting tools: does Turnitin detect AI humanizers.
What This Means for You
To summarize what we have covered:
- Turnitin publishes a list of AI writing models it can detect content from, including GPT, Gemini, Claude, and others.
- The official verb is "can detect content from," and the list includes tools based on these LLMs as well.
- The FAQ promises to expand detection to other models in the future.
- Detection capabilities change over time as the model is updated to target newer LLMs.
- Already-graded documents are not automatically re-evaluated when the model changes.
- The detector trades some missed AI text for a low false positive rate.
If you receive a Turnitin AI report and want to address the flagged passages, import the report and work on them. Eligible passages can be re-run at no charge.
KEEP READING