How to read a Turnitin AI writing report

The percentage is not the share of your document that is AI. It is the share of one specific subset of it - and knowing which subset explains most of what looks wrong in these reports.

HumanPen Team

· 7 min read

The percentage is not a percentage of your document

This is the most misread number in the report. Turnitin defines it as the share of qualifying text, and qualifying text means prose sentences in a long-form writing format: sentences inside paragraphs, in something shaped like an essay, a dissertation, an article.

Everything else is excluded from the calculation. Turnitin states the model does not reliably detect AI writing in non-prose, naming poetry, scripts and code, along with short-form or unconventional writing: bullet points, tables, annotated bibliographies.

A 30 percent score on a paper that is half tables and lists is not "30 percent of your paper". It is 30 percent of the prose Turnitin was willing to assess, which may be a third of what you submitted.

There is also a floor. A submission needs at least 300 words of prose in long-form format before a report is generated at all. Below that there is no percentage - not a low one, none. There is a ceiling too: the AI report caps at 30,000 words, so a full-length thesis does not get one report covering all of it.

The current file requirements add more boundaries: the file must be under 100 MB, be a supported DOCX, PDF, TXT or RTF, and be written in one of the supported languages. Turnitin added Modern Standard Arabic on 18 August 2026, taking that list to four; two of its own English pages still print three, so a page you land on may not have caught up. A successful Similarity Report does not automatically mean the submission qualifies for an AI Writing Report; the two systems have different purposes and requirements.

A denominator example makes the distinction concrete. Suppose a 4,000-word lab report contains 1,600 words in tables, bullet lists, code and references, leaving 2,400 words of qualifying prose. A displayed 25% refers to roughly one quarter of those 2,400 words as classified by the model, not one quarter of all 4,000 words. The report does not show enough information to reconstruct an exact word count by eye, so treat that arithmetic as an explanation of the denominator, not a reverse-engineered total.

Before interpreting the percentage, ask two questions: did the file qualify, and what portion of it counted as qualifying prose?

Why the highlights and the number disagree

Readers routinely notice that the highlighted passages do not look like they add up to the stated percentage and conclude something is broken. Nothing is broken. Turnitin says it directly: a document containing several different writing types will produce a disparity between the percentage and the highlights.

The percentage is computed over qualifying text. The highlights sit in a document that also contains non-qualifying text. The denominators differ, so the two will not reconcile by eye, and checking one against the other is wasted effort.

What the highlights are good for is distribution. Where a flag falls tells you more than how much of it there is.

Read the report in three passes. First, identify the indicator state. Turnitin documents seven of them, and only two put a percentage on the screen: a 0%, and a number from 20 upward. An asterisk, a still-loading report, a processing error, a file that did not meet the requirements, and detection that was switched off at submission time all give you something other than a figure. Second, read the submission breakdown. Until 4 August 2026 it separated AI-generated text from AI-generated text that had been paraphrased; Turnitin has since merged those into one blue category, so reports made before and after that date do not carry the same legend. Third, inspect the highlighted passages in context, including the unhighlighted sentences before and after them. Jumping straight to one bright paragraph removes the context needed to assess it.

  • A cluster in methods or boilerplate may reflect highly conventional language and deserves comparison with required templates or protocols.
  • A boundary between differently drafted passages deserves review of the surrounding sentences, because classification is not necessarily precise at the edge.
  • A scattered pattern should still be checked against version history and sources; visual spread is not a substitute for provenance.
  • No highlights with an asterisk is expected behaviour for the hidden low range, not evidence that the interface failed to load.

Do not count highlighted lines and divide by page length. Page layout, font size, references and excluded writing types make that calculation unrelated to Turnitin's qualifying-text denominator. Use the highlights to locate passages for review, not to audit the percentage as if it were a page-area measurement.

Two categories, and what the second one implies

The report splits its percentage into two kinds of detection.

  • AI-generated only - text the model judges likely to have come from a large language model, possibly then modified by a bypasser.
  • AI-generated text that was AI-paraphrased - text judged likely AI-generated and then likely run through what Turnitin calls "an AI paraphrasing tool or AI word spinner". Turnitin names QuillBot as an example.

The second category is worth sitting with. The paraphraser step is not invisible to the model: running AI text through a rewriter is a pattern the detector was built to recognise as such, rather than a way around it. If you have wondered why paraphrasing did not lower a score, this is the documented reason.

It also explains a common disappointment. The model reads the sentence skeleton, so an edit that leaves that skeleton intact leaves the thing being read intact - a point worth reading alongside what detectors actually measure.

The labels describe the model's classification, not a verified chain of custody. "AI-generated only" does not identify a model, prompt or person; "AI-paraphrased" does not prove that a named service was used. Turnitin itself says its model may misidentify human-written, AI-generated and AI-paraphrased text. The categories help an educator organise review, but they do not convert a probabilistic output into forensic provenance.

There is also a language limitation that is easy to miss: Turnitin's current guide says AI paraphrasing and bypasser detection are included only in the English detector. Spanish and Japanese reports do not currently include those capabilities. A screenshot of the category breakdown therefore needs its language context before it can be compared with another report.

If you see an asterisk instead of a number

Since 8 July 2024, when the model places a result in the 1-19% range, the report shows an asterisk and attributes no numerical percentage and no highlights. A result with no qualifying text is a different state, as is a processing error.

Turnitin's stated reason is that testing found a higher incidence of false positives at low percentages and that hiding the specific number reduces the chance of misinterpretation. The asterisk therefore means "a non-zero result below the reporting threshold, with reduced reliability," not "0%," "safe" or "misconduct."

The rule is about what the interface surfaces; it is not an institutional misconduct threshold. A numerical 19% on a report generated before 8 July 2024 and an asterisk on a newer report may represent the same score band under different display rules. Record the report date before comparing them.

A visible 0% is also not proof that no AI-assisted process occurred. It means the model did not identify qualifying text as likely belonging to its detection categories under that version and configuration. Detection has false negatives as well as false positives, and unsupported or excluded content is outside the percentage altogether.

What the report is not

Turnitin states that its model may misidentify human-written, AI-generated and AI-paraphrased text and should not be used as the sole basis for adverse action against a student. It provides data for educators to evaluate alongside institutional policy and other evidence. The number is an input to human judgement, and the company says so in the report guide.

That is the most useful thing to know if you are on the receiving end of one. The question is not "how do I get this number down" but "what is this number being used for, and what does my institution’s policy say it means".

  • It is not the Similarity Score. Text matching existing sources and text classified as AI are separate analyses; one does not explain the other.
  • It is not a plagiarism finding. Authorship, attribution and permitted assistance are questions for the assessment rules and evidence.
  • It is not a word-by-word probability map. Highlights identify passages assigned to categories; they do not give each word an independent confidence value.
  • It is not stable across all versions. Eligibility, model behaviour and display rules can change, so retain the original report rather than recreating it later and assuming equivalence.

If you need to discuss a report, bring the original submission, the full report, the assignment's AI-use rules and process evidence such as drafts or notes. That set answers four separate questions: what was submitted, what the tool labelled, what assistance was allowed and how the document was produced. The percentage alone answers only the second, and only within the detector's limits.

KEEP READING