Sapling AI Detector vs Turnitin: What the Documentation Actually Says

If you run the same text through Sapling's AI detector and Turnitin, the scores will often differ. This is expected. The two tools use different models trained on different data with different scoring conventions. Here is what Turnitin's documentation says about how its detector works and why the scores will not align.

HumanPen Team

· 14 min read

Sapling and Turnitin Are Separate Systems

Sapling's AI detector and Turnitin's AI writing detection are built by different teams using different underlying models. Each was trained on its own data and classifies text using its own approach. When two detectors use different models, they produce different scores on the same text. A passage that Sapling flags as 70% AI might receive a 25% score from Turnitin, or the reverse. Neither result is necessarily incorrect. They are probability estimates from different models trained on different datasets.

There is no universal AI-detection score that all tools converge on, and why the same text scores differently on every detector goes into where the divergence comes from. Each detector produces its own estimate. Treating scores from different tools as if they measure the same thing in the same units leads to confusion.

We are not making a claim about which tool is more accurate. We are describing what Turnitin's documentation says about its own system and why that system produces scores that will not match Sapling's.

How Turnitin's Scoring Process Works

Turnitin documents a specific process for arriving at its AI writing score:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

The process involves extracting sentences, segmenting them into overlapping units, scoring each segment, pooling sentence scores, and aggregating into a document-level percentage. Other detectors, including Sapling, may use different segmentation strategies, different prediction units, or aggregation methods. The two tools will not produce matching numbers on the same text.

Turnitin's documentation also clarifies that the AI writing detection percentage and the Similarity score are separate. "The Similarity score and the AI writing detection percentage are completely independent and do not influence each other." The Similarity score indicates the percentage of matching text found in the submitted document when compared to Turnitin's comprehensive collection of content for similarity checking. A document can have high similarity and low AI, or vice versa. Turnitin AI writing report vs Similarity report reads the two reports against each other. The two numbers do not influence each other.

Turnitin's Stated False Positive Rate

Turnitin publishes a specific false positive target:

"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."

The documentation restates this as: "In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents." This restatement drops the "over 20% AI writing" qualifier. The version with the 20% threshold is more precise.

To validate this target, Turnitin describes its testing framework: "To bolster our testing framework and diagnose statistical trends of false positives, before every update or new model release, we perform tests on over 700,000 additional academic papers that were written before the release of ChatGPT to further validate our less than 1% false positive rate."

This is what Turnitin says about its own system. For how one percent lands once a university is submitting at scale, see what a 1% false positive rate means. Sapling may have different accuracy targets, different validation processes, or may not publish comparable figures. We cannot make a direct accuracy comparison. We can only report what Turnitin's documentation states.

Short Document Behavior

Document length affects how Turnitin's detector processes text. For short documents, the scoring mechanism behaves differently. The documentation states:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

The overlap mechanism that averages out scores in longer documents does not operate when the text is a single segment. A few hundred words of mixed human and AI content could return a 100% score, which is one of the routes into why Turnitin says your work is 100% AI. This is a structural behavior on short inputs.

If Sapling scores the same short document differently, this is partly because the two tools do not process short inputs the same way.

How Low Scores Are Displayed

Turnitin suppresses numeric scores below 20%. The documentation states:

"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."

A document that Turnitin's model estimates at 8% AI will show an asterisk in the report, not the number 8%. No highlights are applied to the text. The rationale is to avoid surfacing potential false positives in a low-confidence range, unpacked further in what the asterisk (*%) means on a Turnitin AI score.

Sapling may display scores in this range as a full percentage. Other tools may highlight text at any non-zero level. Turnitin deliberately suppresses both the number and the highlights below 20%. This scoring convention means the two tools will present very different outputs for the same low-AI document.

What This Means for You

If you are comparing Sapling and Turnitin results on the same document, here is the summary:

  • Separate models, separate scores. Sapling and Turnitin use different detection models. Their scores are not expected to match.
  • Turnitin scores by segmenting and aggregating. Sentences are extracted, segmented into overlapping units, scored, pooled, and aggregated. The AI score and Similarity score are fully independent.
  • Turnitin targets under 1% false positives for documents with over 20% AI writing, validated against over 700,000 additional academic papers written before ChatGPT.
  • Short documents get all-or-nothing scores. A few hundred words may come back as 0% or 100% even for mixed text.
  • Scores below 20% show as an asterisk. No numeric percentage or highlights appear in the 1-19% range.

Eligible passages can be re-run at no charge.

KEEP READING