Originality.ai vs Turnitin AI Detection: What the Documentation Says

Students sometimes run their paper through Originality.ai and get a different score than Turnitin. This is expected. Different detectors use different models trained on different data. Turnitin's FAQ describes its mechanism, false positive targets, and the asterisk convention for low scores. Here is what the documentation says.

HumanPen Team

· 10 min read

The Short Answer

Originality.ai and Turnitin are different AI detectors built on different models with different training data and different classification thresholds. A score from one will not predict the score from the other. Turnitin's FAQ describes its mechanism: text is split into overlapping segments, each classified with a probability score between 0 and 1. The FAQ also states a false positive target of under 1% for documents with over 20% AI writing, and uses an asterisk convention for scores in the 1-19% range. Originality.ai publishes its own accuracy claims, but those are vendor-reported and cannot be independently verified against Turnitin's documentation. The score that matters is the one from the tool your institution uses. If your school uses Turnitin, no third-party score can tell you what your Turnitin score will be.

How Turnitin's Detector Works

The FAQ describes the detection process:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."

Segments overlap, meaning sentences can receive multiple scores that are pooled together. This pooling smooths individual prediction errors and contributes to the overall document score. Originality.ai uses its own classification approach, which is not documented in Turnitin's materials. The two systems are built differently and evaluate text differently.

Turnitin's False Positive Target

The FAQ states a specific target:

"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."

The next sentence: "In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents."

This target is specific to Turnitin and applies to documents with over 20% AI writing. Originality.ai may have a different false positive rate, but comparing the two numbers requires understanding how each vendor defines and measures false positives. We do not make accuracy claims about third-party detectors. What the documentation tells us is that Turnitin has explicitly stated this target for its own model. What that target costs at scale is worked out in what a 1% false positive rate means when a university submits 75,000 papers.

The Asterisk Convention

Turnitin has a specific display rule for low scores:

"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."

This means a Turnitin score showing "%" indicates AI was detected below 20%, but the exact percentage is withheld to avoid overstating accuracy in a range prone to false positives. Originality.ai does not use this convention and will display a specific percentage in this range. A paper showing "%" on Turnitin might show "15%" on Originality.ai. This does not mean one tool is more accurate than the other. It means they handle the low-confidence range differently. What the asterisk (*%) means on a Turnitin AI score covers the Turnitin side of it.

Why Scores Disagree Between Tools

Different AI detectors produce different scores on the same text because they use different models, different training data, and different thresholds. Turnitin's FAQ describes how its model pools overlapping segment scores into a document-level percentage. A different detector might use a different segmentation strategy, a different scoring scale, or a different aggregation method. The FAQ also notes that short documents produce all-or-nothing predictions: "In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap." Originality.ai may or may not have the same behavior on short documents.

The AI score and the similarity score are also independent on Turnitin: "The Similarity score and the AI writing detection percentage are completely independent and do not influence each other." A detector that combines AI and similarity into a single score will produce a number that is not comparable to Turnitin's separate AI percentage. Why the same text scores differently on every detector collects the reasons.

What This Means for You

To summarize what we have covered:

  • Originality.ai and Turnitin are different detectors with different models. Scores will not match.
  • Turnitin's detector splits text into overlapping segments and assigns probability scores between 0 and 1.
  • Turnitin's false positive target is under 1% for documents with over 20% AI writing.
  • Turnitin uses an asterisk for scores in the 1-19% range to avoid overstating accuracy.
  • Short documents get all-or-nothing predictions on Turnitin due to single-segment evaluation.
  • The score that matters is the one from the tool your institution uses.

If you receive a Turnitin AI report and want to address the flagged passages, import the report and work on them. Eligible passages can be re-run at no charge.

KEEP READING