Does Turnitin Detect LLaMA? What the Model List Says
LLaMA appears on Turnitin's published list of detectable models. We examine what the list says, how the detection mechanism works, why some AI text goes undetected, and what you should do next if your report comes back flagged.
HumanPen Team
· 10 min read
The short answer
Yes. LLaMA is explicitly named on the model list that Turnitin publishes. The list states that Turnitin's AI writing detection model for English submissions can detect content from a set of models that includes LLaMA, along with GPT, Gemini, Claude, Mistral, DeepSeek, Nova, Grok, and o1-mini. The list also covers "tools based on these LLMs as well," which means applications and services built on top of LLaMA's language models fall within the stated detection scope.
LLaMA is an open-source model family, which adds a wrinkle. Many fine-tuned variants and derivatives exist, built by different organizations on top of the base LLaMA models. The "tools based on these LLMs" clause in Turnitin's wording is meant to cover exactly these cases; the same open-weights wrinkle comes up in does Turnitin detect Mistral AI text. Whether every fine-tuned variant is caught with the same accuracy is a separate question, one the list itself does not answer.
What the model list actually says
Turnitin does not publish a table of models with pass/fail detection results. Instead, it provides an inline comma-separated list within its documentation. The introductory sentence reads: "Currently, Turnitin's AI writing detection model for English submissions can detect content from" a list that includes LLaMA alongside other major language models. Claude and Gemini are on that same line, and does Turnitin detect Claude, Copilot or Gemini reads the closed-weights half of it. The list concludes by stating coverage extends to "tools based on these LLMs as well."
The key phrase here is "can detect content from." This wording tells us the detector is trained to recognize text generated by these models. It does not promise perfect detection. The sentence that immediately follows the list is equally important: "We will continue to expand our detection capabilities to other models in the future."
For LLaMA specifically, the open-source nature of the model family means there are many downstream tools and fine-tuned variants. The "tools based on these LLMs" clause is designed to cover these derivatives, bringing them within the stated detection scope. The sensitivity of that coverage can vary, and we address that in the sections below.
How the detection mechanism works
To understand why detection is not a binary yes or no, it helps to look at the mechanism described by Turnitin. When a paper is submitted, sentences are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score.
This means the final percentage on your report is an aggregation of many small, sentence-level predictions. It is not a single scan that stamps the whole document as AI or human. Each sentence contributes to the total, and overlapping segments help smooth out borderline cases.
Detection capabilities change over time
Turnitin's model is not static. According to the documentation, "As we iterate and develop our model further to better detect newer LLMs, it is likely that our detection capabilities will also change, affecting the AI percentage.. However, for a submitted document, the AI percentage will change only if it's re-submitted again to be processed."
This has a practical implication. If you submit a document today and receive a certain AI percentage, that same document could receive a different percentage if re-submitted weeks or months later. The model updates, and those updates can shift scores in either direction. A passage that reads as human today might read as AI-generated after an update, or vice versa.
For LLaMA text, this means the detection rate you experience now may not match what you experience later. The model list confirms LLaMA is in scope, but the sensitivity of that detection can vary as Turnitin refines its model.
Why some AI text gets missed (and what drives detection)
Turnitin is transparent about the fact that its detector does not catch everything. The documentation states: "In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document. We're comfortable with that since we do not want to incorrectly highlight human-written text as AI-written. For example, if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."
This is a deliberate trade-off. Turnitin prioritizes a low false-positive rate, which means it accepts that some AI text will slip through. A document showing 50% AI could actually contain up to 65% AI content. The detector errs on the side of caution, preferring to miss AI text rather than flag human text incorrectly.
Understanding why the detector works this way requires looking at what it actually measures. Turnitin's model does not rely on the metrics often discussed in public conversations about AI detection, which is the whole subject of does Turnitin use perplexity and burstiness to detect AI. The documentation states: "Our model is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions. Instead, it learns statistical patterns from our training data." The model then applies what it has learned to classify text. As Turnitin puts it, "Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers."
In other words, the detector looks at word probability patterns, not surface-level signals like sentence length variation or vocabulary rarity. This is why heavily edited or rewritten AI text can sometimes evade detection, and why does Turnitin detect AI humanizers turns on how far the rewrite went. If the word probability sequences shift enough during editing, the classifier may no longer recognize the text as AI-generated. For LLaMA variants, this means the detector looks at the statistical patterns in the text itself, not at which specific variant produced it.
What to do if your report is flagged
If your Turnitin report flags passages as AI-generated, the most practical step is to address only the flagged sections rather than rewriting the entire document, the approach in rewriting only the paragraphs a Turnitin report flagged. Turnitin's sentence-level scoring means you can identify specific segments that triggered the detection and focus your revision there.
Keep your citations, tables, and formatting intact. Only the passages flagged as AI need attention. If you revise and re-submit, remember that the AI percentage may shift due to model updates, not just because of your edits.
Eligible passages can be re-run at no charge, so you can refine flagged sections without additional cost until the percentage reaches an acceptable level.
KEEP READING