Winston AI vs Turnitin AI Detection: Why the Scores Differ
Students sometimes run their paper through Winston AI before submitting to Turnitin. The two tools use different models and produce different scores. Turnitin's FAQ describes its mechanism, false positive targets, and the asterisk convention. Here is why cross-tool comparisons fail.
HumanPen Team
· 11 min read
The Short Answer
Winston AI and Turnitin are different AI detectors built on different models with different training data and different classification thresholds. A score from one will not predict the score from the other. Turnitin's FAQ describes its mechanism: text is split into overlapping segments, each classified with a probability score between 0 and 1. The FAQ also states a false positive target of under 1% for documents with over 20% AI writing, and uses an asterisk convention for scores in the 1-19% range. Winston AI publishes its own accuracy claims, but those are vendor-reported and cannot be verified against Turnitin's documentation. The score that matters is the one from the tool your institution uses. If your school uses Turnitin, no third-party score can tell you what your Turnitin score will be.
How Turnitin's Detector Works
The FAQ describes the detection process:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."
Segments overlap, meaning sentences can receive multiple scores that are pooled together into a document-level percentage. Winston AI uses its own classification approach, which is not documented in Turnitin's materials. The two systems are built differently and evaluate text differently.
Turnitin's False Positive Target
The FAQ states a specific target:
"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."
The next sentence: "In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents."
This target is specific to Turnitin and applies to documents with over 20% AI writing. Winston AI may report a different false positive rate, but comparing the two numbers requires understanding how each vendor defines and measures false positives. We do not make accuracy claims about third-party detectors.
The Asterisk Convention
Turnitin has a specific display rule for low scores:
"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."
Winston AI does not use this convention and will display a specific percentage in this range. A paper showing "%" on Turnitin might show "14%" on Winston AI. This does not mean one tool is more accurate than the other. It means they handle the low-confidence range differently. [What the asterisk (%) means on a Turnitin AI score](/blog/what-does-the-asterisk-mean-on-turnitin-ai-score) covers the Turnitin side of it.
Short Documents and All-or-Nothing Predictions
The FAQ describes how short documents behave:
"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."
The next sentence: "This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."
This is specific to Turnitin's segmentation approach. Winston AI may or may not exhibit the same behavior on short texts. If you test a short excerpt on Winston AI and get 0%, that tells you nothing about what Turnitin will do. The far end of that behaviour is Turnitin saying your work is 100% AI.
The Score That Matters
The AI score and the similarity score are independent on Turnitin: "The Similarity score and the AI writing detection percentage are completely independent and do not influence each other." Some detectors combine AI and similarity into a single number, which makes their scores structurally non-comparable to Turnitin's separate AI percentage. The only score that matters for your submission is the one produced by the tool your institution uses. Why the same text scores differently on every detector explains why the spread is structural.
What This Means for You
To summarize what we have covered:
- Winston AI and Turnitin are different detectors with different models. Scores will not match.
- Turnitin's detector splits text into overlapping segments and assigns probability scores.
- Turnitin's false positive target is under 1% for documents with over 20% AI writing.
- Turnitin uses an asterisk for 1-19% scores to avoid overstating accuracy.
- Short documents get all-or-nothing predictions on Turnitin.
- The only score that matters is the one from the tool your institution uses.
If you receive a Turnitin AI report and want to address the flagged passages, import the report and work on them. Eligible passages can be re-run at no charge.
KEEP READING