How Wide Is the Margin on a Turnitin AI Score? We Counted
If a paper comes back at 43%, the natural question is 43 plus or minus what. The three vendor pages that define the report do not publish an answer. The gap is not limited to one FAQ: it appears across all three pages we checked. What those pages publish instead is one directional example, one suppressed band, and a mechanism description that tells you where the real variation lives.
HumanPen Team
· 13 min read
What is the margin of error on a Turnitin AI score?
There is not a published one. Across the three pages that define the AI writing report, read on 28 August 2026 and totalling 46,712 characters of rendered article text, the words `interval`, `margin`, `standard deviation`, `uncertain`, `precision`, `variance`, `error bar`, `plus or minus`, `±` and `estimat` return zero hits between them. The counter was working: `percentage` returns 47 on the same corpus and `false positive` returns 21. The vendor does publish three numeric things that people mistake for a margin, and none of them is one.
That absence is the useful finding, not a complaint. It tells you which question to stop asking and which two questions are actually answerable about the number on your report.
Two neighbouring pages here cover different ground and this one does not repeat them: Turnitin's False Positive Rate, Explained is about the error rate of the yes/no decision, and What a 1% false positive rate means when a university submits 75,000 papers is about what that rate does at institutional scale. This page is about the number itself: what it is an estimate of, and how far it could move.
The count, with its denominator
Three pages, read logged out in a browser on 28 August 2026, measured as `document.querySelector('article').innerText`, not the page body. The body adds roughly 725 characters of navigation and would inflate the denominator without adding any relevant text. Character counts: Using the AI Writing Report 6,492; the AI writing detection capabilities FAQs 30,987, which the page itself marked "Updated 9 days ago"; the AI writing detection model release notes 9,233. Total 46,712.
| Search term | Report page | FAQs | Release notes | Total |
|---|---|---|---|---|
| `interval` | 0 | 0 | 0 | 0 |
| `margin` | 0 | 0 | 0 | 0 |
| `standard deviation` | 0 | 0 | 0 | 0 |
| `uncertain` | 0 | 0 | 0 | 0 |
| `precision` | 0 | 0 | 0 | 0 |
| `variance` | 0 | 0 | 0 | 0 |
| `error bar` | 0 | 0 | 0 | 0 |
| `plus or minus` and `±` | 0 | 0 | 0 | 0 |
| `estimat` | 0 | 0 | 0 | 0 |
| `percentage` (control) | 10 | 31 | 6 | 47 |
| `false positive` (control) | 2 | 12 | 7 | 21 |
| `qualifying text` (control) | 6 | 4 | 3 | 13 |
| `segment` (control) | 0 | 9 | 1 | 10 |
| `reliab` (control) | 2 | 3 | 0 | 5 |
| `probabilit` (control) | 0 | 4 | 0 | 4 |
The controls are there because a row of zeros is worthless without them. If the search had been broken, or the page had failed to render, `percentage` would have come back zero too. It came back 47.
Two near-misses worth printing, because they are the kind of thing a careless read turns into a finding. `confidence` returns exactly one hit in the whole corpus, and it is not statistical: "While Turnitin has confidence in its model, Turnitin does not make a determination of misconduct." And `precise` returns exactly one, in a 2023 release note about segment boundary detection. Neither is a dispersion measure. If you repeat this count and land on either of them, read the sentence before treating it as one.
And the statement runs both ways. "The vendor has not published a margin" is a fact about these pages on this date. It does not mean the model has no internal measure of its own confidence; the mechanism section below says it computes one. It equally cannot be used by anyone to argue the number is precise. It means the question is unanswered in the pages checked, which is different from being answered either way.
The three numbers people mistake for a margin
The documentation is not silent about error. It is silent about this kind of error. Three published numbers get pressed into service as a margin and none of them fits.
- The under-1% false positive target. Its stated scope is "documents with over 20% of AI writing", and it is about how often a fully human document gets flagged at all. That is the error rate of a decision across a population, not a band around your 43.
- The "50% could be 65%" example. This one is closest, and it is the subject of the next section. It is directional and it is about missed AI text, not about the displayed figure being too high.
- The suppressed 1–19% band. No number is displayed below the 20% threshold. That is a statement about where the vendor considers the reading unsafe to show, which is informative, but a suppression rule is not a margin either.
Missing from that list is the thing you would need: any statement of the form "a document scored X will typically fall between Y and Z". Nothing in 46,712 characters takes that shape. So "43% plus or minus what" has no answer backed by these primary sources. Anyone offering a range should be able to point to a different primary source for it.
The direction in the closest published example
The closest passage we found to a margin of error quantifies how far the reported figure can sit from reality, and it points one way only.
"In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document. We're comfortable with that since we do not want to incorrectly highlight human-written text as AI-written. For example, if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."
Read that carefully, because the shape matters more than the numbers. It says the true amount of AI text can be higher than the figure shown. It does not say the reverse. Our count found no companion range in the other direction on those three pages. The only `up to` hit in the release notes refers to a 30,000-word input limit, not a percentage.
That is the full extent of the asymmetry this passage supports. It explains why the example points upward: the system accepts missing some AI text in order to reduce false positives. It does not tell us how much evidential weight to give an individual high or low score. The checked pages publish no calibration curve, likelihood ratio or base-rate-adjusted measure for that inference, and the same FAQ says the percentage should not be the sole basis for action.
That paragraph offers no support when you wrote the paper yourself and believe the score is wrong. It is not a margin; it is a disclosure about which way the system errs. The material that does help with a false positive is a different subject: Turnitin's False Positive Rate, Explained covers the vendor's own statements on that.
Where the variation actually lives
The FAQs lay out the scoring sequence. Once you have read it, you can see where an uncertainty measure would have gone and where it drops out of the public description.
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."
So there are four steps, and a continuous quantity exists at step one. Each segment gets a value between 0 and 1. Sentences inherit it. Overlapping scores are pooled. The remaining sentence-level results are combined into one figure for the document. A probability of 0.51 and a probability of 0.99 are very different states of knowledge, and by the time they reach your report they have both become part of one integer percentage.
The three pages do not state the pooling rule, the aggregation rule, or the threshold at which a pooled sentence score becomes a flagged sentence. That leaves three unstated functions between the model's uncertainty and the percentage shown to your instructor. It is a more exact account of the public gap than simply reporting a row of zeros.
None of that makes the figure meaningless. It is a model-derived estimate of the proportion of qualifying text that is likely AI-generated. The missing pooling rule, aggregation rule and margin limit what can be inferred from that estimate; they do not erase its published unit. The careful reading of 43% is therefore "the model estimated 43% of the qualifying text", not "43% of every word in this document, with a known error band".
The documented short-document exception
The documentation supports one narrow warning about length, not a general rule that predicts score resolution from a document's word count.
The FAQs say: "In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap. This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated." That is a documented single-segment case. The checked pages do not publish a general function showing how score resolution changes as documents get longer.
So the practical question to add is not an invented margin but the report's context: how much qualifying text entered the estimate, and is this the few-hundred-word single-segment case the vendor specifically warns about? Do not extrapolate that short-document behavior into an unpublished rule for an 8,000-word chapter. Short Documents and the All-or-Nothing Problem in Turnitin works through the documented short case, and What Is "Qualifying Text" in Turnitin AI Detection? covers what does and does not enter the denominator in the first place.
The band where the vendor stops showing a number
There is one place where Turnitin has effectively said the reading is not solid enough to print, and it did so by removing the number.
"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."
Two things follow that are worth carrying. First, an asterisk is not a low score being hidden from you; it is the vendor declining to attribute a score, and there are no highlights either, so there is nothing specific for anyone to point at. Second, this is the clearest place in the three pages where the vendor withholds part of its output because of false-positive risk. It is the nearest thing we found to an uncertainty disclosure, and it takes the form of a rule rather than a number. What Does the Asterisk (*%) Mean on a Turnitin AI Score? goes through what that display actually means.
How to read a number that ships without one
Given all of the above, here is what survives as usable when someone puts a percentage in front of you.
- Stop asking for the margin and ask for the denominator. How many qualifying sentences did this figure come from? A few hundred words means near-binary output by the vendor's own description.
- Note the direction of the published example. Reported 50 could be 65. We found no source-backed range in the other direction on the three pages checked.
- Treat the asterisk as an absence of attribution, not a small number. No score, no highlights.
- Do not convert the estimated share of qualifying text into a share of your paragraphs. It comes from unpublished pooling and aggregation over segments, not a count of sentences you can locate one to one.
- Do not assign significance to a cross-draft change that the vendor has not defined. The checked pages publish no significance threshold for a change of a few points. Keep each report's qualifying-text context and highlights with the number.
- Use the vendor's own framing when it is quoted at you. Its FAQ says the percentage "should not be used as the sole basis for action or a definitive grading measure by instructors", and that its model "may not always be accurate".
If someone hands you a confidence interval for a Turnitin AI score, ask where it came from. On the evidence above, it did not come from Turnitin.
If specific passages are going to be rewritten anyway
Everything above is about reading the report rather than acting on it, and the two need to stay separate. But there is one practical consequence worth stating, because it follows directly from the counting problem: if a percentage has no published uncertainty attached to it, then chasing the percentage is chasing something with no defined target. The highlighted passages do have a defined target. They are specific spans of text you can look at.
That is the whole reason HumanPen, which we build, is organised around importing the report rather than around the number on it. Read this as a first-party description. Hand over the AI writing report you already have and the flagged spans become the boundary of the job; you approve that boundary on screen, only the confirmed paragraphs are rewritten, and everything else is left unchanged. That is a rewrite-scope statement, not a claim about upload, transport or storage. The reason to prefer that over a whole-document pass is not a better figure, which nobody can promise. Each changed paragraph adds another source check before submission, so narrow scope keeps that verification queue short. Where the file itself is the deliverable, What happens to citations, tables and equations in an AI humanizer? covers what to check afterwards.
And the obvious caveat, which this page has been building towards for eight sections: nobody can tell you what the next figure will be, us included. If you wrote the work yourself, none of this applies. The argument to make is about evidence and process, not about the text.
Frequently asked questions
Does Turnitin publish a margin of error for the AI percentage? Not on the three pages that define the report. Across 46,712 characters read on 28 August 2026, `interval`, `margin`, `standard deviation`, `uncertain`, `precision`, `variance`, `error bar`, `plus or minus`, `±` and `estimat` all return zero, while `percentage` returns 47 and `false positive` returns 21 on the same text.
Is 43% really 43% of my paper? It is the model-derived estimate that 43% of the qualifying text is likely AI-generated, not necessarily 43% of the whole submission. Because the pooling and aggregation rules and a margin are not published, you cannot map it one to one onto visible paragraphs.
Could the real amount of AI text be higher than the number shown? The vendor says so directly, with its own example: identify 50% and the document "could contain as much as 65% AI writing". That example points upward. We found no source-backed range in the opposite direction on the three pages checked.
Why does a short essay get an extreme score? Because there may only be one segment. The FAQs say that in documents of a few hundred words the prediction is "mostly 'all or nothing'", and that mixed text can therefore be flagged as entirely AI-generated.
Why does my report show an asterisk instead of a number? Below the 20% threshold the vendor attributes no score and no highlights, and displays `*%` instead, stating that this is to avoid potential incidence of false positives.
Does a low or high score prove who wrote my document? No. The vendor's example says a displayed share may understate AI text, but the checked pages do not publish the calibration and base-rate information needed to turn an individual percentage into evidential weight. The vendor also says the percentage should not be the sole basis for action.
KEEP READING