What Is "Qualifying Text" in Turnitin AI Detection? A Detailed Explanation
When Turnitin reports an AI writing percentage, that number does not cover your entire submission. The model only analyzes "qualifying text," which means prose sentences in standard grammatical form. Lists, bullet points, and non-sentence structures are excluded. This explains why the percentage and the highlighted passages sometimes do not match. Here is what the documentation says about what qualifies, what does not, and why it matters.
HumanPen Team
· 15 min read
What "Qualifying Text" Means
Turnitin's AI detection model does not analyze every word in a document. It operates on a specific subset of the text called "qualifying text." The documentation defines this clearly.
"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."
So the model looks for prose: blocks of text organized into standard grammatical sentences. Anything that does not meet this bar is left out of the analysis. Bullet points, numbered lists, table headings that are not full sentences, and other non-prose structures are excluded from the AI detection scope.
The model's own documentation acknowledges this limitation: "The model does not reliably detect AI-generated text in the form of non-prose, or code, nor does it detect short-form/unconventional writing such as bullet points (short non-sentence structures)."
If a document is heavy on lists and light on prose, the AI detection model has less material to work with, and the reported score reflects only the prose portions. There is a floor under that as well, since the file requirements ask for at least 300 words of prose before a report is generated at all, which can Turnitin check a whole thesis goes into.
The Disparity Between Percentage and Highlights
One of the most confusing aspects of Turnitin's AI detection report is that the percentage and the highlighted passages can appear to tell different stories. The documentation explains why.
"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."
"This percentage is not necessarily the percentage of the entire submission."
If your document is 50% prose and 50% bullet points, and the AI detection model flags 40% of the prose as AI-generated, the report might show a 40% AI score. But that 40% refers to the qualifying prose, not the entire document. Someone reading the report without understanding this could assume the entire submission is 40% AI-generated, when the actual proportion of the full document is lower.
This disparity is not a bug. It is a consequence of the model's scope. Understanding it helps you read the report accurately, and how to read a Turnitin AI writing report works through the rest of the page.
How the Model Processes Text
To understand why qualifying text matters, it helps to know what the model does with it. The documentation describes the mechanism:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."
The model extracts sentences, segments them into overlapping sections, and classifies each segment. Each segment gets a probability score between 0 and 1. Only qualifying prose sentences go through this pipeline. Non-prose content is never segmented or scored, because the model does not treat it as analyzable text.
This is why the scope of "qualifying text" directly determines what the percentage represents. The score is an aggregate of segment-level probability scores, and only qualifying prose sentences are part of that aggregate.
Tables Are In, Bibliographies Are Out
Two specific content types have their own rules.
Tables: "We are now able to process long-form prose text in tables." So prose inside table cells, if it is written in standard grammatical sentences, is now part of the qualifying text. Resubmit to reprocess existing submissions that contain tables. If you have a document with meaningful prose in tables and it was submitted before this update, the AI score might not reflect the table content unless you resubmit.
Bibliographies: "We have fixed a bug that was occasionally highlighting AI writing within references listed in a bibliography. Bibliographies are now excluded when processing the AI writing report." Resubmit to reprocess existing submissions that contain highlighted reference sections. If your previous report showed flagged text inside your reference list, that was a bug, and resubmitting will exclude those sections from analysis. Why is Turnitin flagging my references and citations covers how far the exclusions reach on the similarity side.
These two rules mean the definition of qualifying text is not static. It has been refined to include table prose and to exclude bibliographies, which is one reason your Turnitin AI score can change on resubmission without the text changing.
False Positives in Qualifying Prose
Even within qualifying prose, certain types of writing are more prone to false positives. The documentation identifies specific patterns.
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."
"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
So if your prose is structurally monotonous, repeats the same phrasing, or is paraphrased without adding new ideas, the model may flag it as AI-generated even when it is human-written. This is not a failure of the qualifying text filter, it is what AI detectors measure doing exactly what it was trained to do. It is a separate issue within the prose that does qualify for analysis.
What This Means for You
If you are trying to understand why your Turnitin AI report looks the way it does, here is the summary:
- Only prose sentences qualify. Lists, bullet points, and non-sentence structures are excluded from analysis. The model focuses on standard grammatical sentences.
- The percentage covers qualifying prose only. It is not the percentage of the entire submission. This creates a disparity between the number and the highlights.
- Table prose is now processed. Long-form prose in tables counts as qualifying text. Resubmit to reprocess existing submissions that contain tables.
- Bibliographies are excluded. A bug that flagged references has been fixed. Resubmit to reprocess existing submissions that contain highlighted reference sections.
- Watch for false positives. Structurally flat, repetitive, or paraphrased prose is more likely to be flagged. Take this into account when reading the percentage.
Eligible passages can be re-run at no charge.
KEEP READING