Copyleaks vs Turnitin AI Detection: Why Different Detectors Disagree

Students often run their paper through multiple AI detectors and get completely different scores. Turnitin might say 20% while Copyleaks says 80%. The reason is simple: different models, different training data, different classification thresholds. Understanding why detectors disagree helps you read each score in context.

HumanPen Team

· 8 min read

The Short Answer

Different AI detectors give different scores because they are different models trained on different data with different classification thresholds. Copyleaks and Turnitin are separate products built by separate teams. Each has its own model, its own training corpus, and its own approach to aggregating segment-level predictions into a document score. When Copyleaks says 80% and Turnitin says 20% on the same text, neither is necessarily wrong. They are measuring different statistical patterns and applying different decision boundaries. The score you should care about is the one from the tool your institution actually uses. If your school uses Turnitin, your Copyleaks score is informative but not authoritative.

How Turnitin's Detector Works

Turnitin's FAQ describes its pipeline:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

Turnitin also publishes a specific false positive target: "We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing." This target shapes how aggressively the model classifies borderline text. A model tuned to minimize false positives will classify more uncertain segments as human, which can produce lower AI scores on mixed text.

Why Copyleaks and Turnitin Disagree

Copyleaks uses its own AI detection model with its own training data and classification thresholds. The details of Copyleaks' model are not public in the same way Turnitin's FAQ is public, so we cannot do a side-by-side comparison of their mechanisms. But the structural reason for disagreement is clear: two different models trained on different data will classify the same text differently. A sentence that Turnitin's model classifies as 0.3 probability of AI might get 0.7 from Copyleaks' model. When those segment-level differences are aggregated across a document, the final percentages can diverge significantly. Why the same text scores differently on every detector takes that apart in general.

There is also a threshold difference. Turnitin does not display scores between 1% and 19%, showing only an asterisk instead. The FAQ explains: "To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range." If Copyleaks displays a 15% score for the same text that Turnitin shows as an asterisk, the discrepancy is partly a display decision, not just a classification difference — see what the asterisk (*%) means on a Turnitin AI score.

Short Documents Amplify Disagreement

Turnitin's FAQ notes that short documents produce less reliable predictions:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

The next sentence: "This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

On a short document, one detector might classify the entire text as AI (getting 100%) while another classifies it as human (getting 0%). The all-or-nothing behavior means that small differences in how each model handles a single segment can produce dramatic differences in the final score. If you are comparing scores across detectors, document length matters. A 300-word abstract might get wildly different scores from different tools, while a 3,000-word paper might show closer agreement.

Which Score Matters

The score that matters is the one from the tool your institution uses. If your university uses Turnitin, your Turnitin AI score is the one that affects your academic standing. A Copyleaks score of 5% does not help you if your Turnitin score is 80%. A Copyleaks score of 80% does not hurt you if your institution does not use Copyleaks. Running your paper through multiple detectors can give you a general sense of whether your text has AI-like patterns, but it cannot tell you what your institution's detector will say. The only way to know your actual score is to see the report from the tool your institution uses, and only instructors and administrators can see the Turnitin AI indicator. Whether your school runs one at all depends on what they bought.

What This Means for You

To summarize what we have covered:

  • Different AI detectors use different models with different training data and thresholds. Disagreement is expected, not surprising.
  • Turnitin publishes its false positive target (under 1% for documents with over 20% AI) and its mechanism (overlapping segment classification).
  • Turnitin suppresses scores between 1% and 19% as asterisks, which can amplify apparent disagreement with other tools.
  • Short documents produce all-or-nothing predictions, which can cause dramatic divergence between detectors.
  • The score that matters is the one from the tool your institution actually uses.

If you have a Turnitin report showing which passages were flagged, import the report and work on those specific passages. Eligible passages can be re-run at no charge.

KEEP READING