Does Turnitin Detect Gemini and Copilot? What the FAQ Says

Gemini and Copilot are two of the most common AI writing tools students ask about. Turnitin's FAQ publishes a list of models it can detect content from, and Gemini is on it. Copilot is built on GPT models, and the list covers tools based on these LLMs as well. The FAQ also says detection capabilities change over time as the model is updated. Here is what the documentation actually says.

HumanPen Team

· 11 min read

The Short Answer

Yes for Gemini, and effectively yes for Copilot. Turnitin's FAQ publishes a list of AI writing models its detector can identify content from. The list includes Gemini, along with GPT, Claude, LLaMA, Mistral, Deepseek, Nova, Grok, and o1-mini. The official verb used is "can detect content from." The list ends with: "and tools based on these LLMs as well." Copilot is built on GPT models, and GPT appears on the list. This means Copilot falls under the category of tools based on a listed LLM. The FAQ adds: "We will continue to expand our detection capabilities to other models in the future." The FAQ also acknowledges that detection capabilities are not static and will change as the model is updated.

What the Model List Says

The FAQ provides a comma-separated list of models it can detect content from. The introductory sentence reads: "Currently, Turnitin's AI writing detection model for English submissions can detect content from …" followed by the list including GPT, Gemini, Claude, LLaMA, Mistral, Deepseek, Nova, Grok, and o1-mini. The list concludes:

"and tools based on these LLMs as well."

The next sentence: "We will continue to expand our detection capabilities to other models in the future."

This phrasing matters. The list does not name every individual tool or wrapper. It names the underlying model families and then covers anything built on top of them. Copilot is powered by GPT models, so it falls within the scope of "tools based on these LLMs." Gemini is named directly on the list. The promise to expand means the list may grow as new models emerge and the detection model is updated.

How the Detector Processes Your Text

The FAQ describes the detection mechanism in terms of segmentation and probability scoring:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."

Each segment gets its own score. The overall percentage reflects the proportion of segments the model classifies as AI-generated. The detector is not looking for a watermark or a single telltale feature. It is classifying text segment by segment based on probability scores. What AI detectors measure beyond perplexity and burstiness covers what goes into each of those scores.

What the Detector Might Miss

The FAQ acknowledges a trade-off between false positives and missed AI text:

"In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document."

The detector is tuned to avoid falsely accusing students of using AI. The cost of that tuning is that some AI-generated text may go undetected. This is a deliberate design choice documented by Turnitin itself. A model that produces text closer to human word probability patterns would be harder to detect, at least until the detection model is updated to cover it. The same lag is the whole question in does Turnitin detect AI humanizers.

Detection Capabilities Change Over Time

The FAQ explicitly states that detection is not a fixed target:

"As we iterate and develop our model further to better detect newer LLMs, it is likely that our detection capabilities will also change, affecting the AI percentage."

The next sentence: "However, for a submitted document, the AI percentage will change only if it's re-submitted again to be processed."

Two implications follow. First, the detection model is being updated to target newer LLMs, which means coverage of models like Gemini and Copilot may improve or shift over time. Second, a document scored before a model update retains its original score unless it is re-submitted. The score on a past submission reflects the detection capabilities at the time of that submission, not the current ones, which is why still flagged after a Turnitin recheck argues for comparing two reports rather than two numbers.

What This Means for You

To summarize:

  • Gemini appears directly on Turnitin's published list of models it can detect content from.
  • Copilot is built on GPT models, which appear on the list. The list also covers "tools based on these LLMs as well."
  • The FAQ promises to "continue to expand our detection capabilities to other models in the future."
  • The detector processes text in overlapping segments and assigns each a probability score between 0 and 1.
  • To maintain a low false positive rate, the detector may miss some AI-generated text.
  • Detection capabilities change over time. A document scored before a model update keeps its original score unless re-submitted.

If you receive a Turnitin AI report and want to address the flagged passages, import the report and work on them. Eligible passages can be re-run at no charge.

KEEP READING