Why did Turnitin flag my references?
Two different reports, two different things being measured, and only one of them can put colour on a bibliography. Here is how to tell which one you are holding, why a sentence with a citation in it can still get flagged, and where Turnitin says its own false positives cluster.
HumanPen Team
· 7 min read
First, work out which report you are looking at
Turnitin excludes bibliographies when it processes the AI writing report, and AI writing highlights never appear in the Similarity Report at all. So a highlighted reference list is a similarity match, not an AI flag, and your bibliography is not what raised your AI percentage.
That one distinction settles most of these questions before you have to think about anything else. Turnitin produces two separate things from a single submission and they measure different properties. The Similarity Report finds text in your submission that matches text elsewhere. The AI Writing Report estimates how much of your prose was generated by AI. Turnitin's documentation describes the two as independent of each other and states that the AI writing highlights do not appear in the Similarity Report. Neither number feeds the other one.
There is a practical reason people end up unsure which one they are holding. Only instructors and administrators can see the AI writing indicators; students cannot. An instructor can download a PDF of the AI report and share it, so what usually reaches a student is that PDF, or a forwarded screenshot with nothing on it that says which report it came from.
Ask which report it was before anything else. Every branch below depends on the answer.
For a fuller walkthrough of how the two differ, see AI Writing Report vs Similarity Report.
If it was the Similarity Report, matched references are the ordinary case
A reference entry is the most repeatable string in academic writing. Author, title, journal, volume, year, and the punctuation between them are all fixed by the style guide, so two papers citing the same source produce the same line, character for character. That is what a citation style is for.
A bibliography lighting up in a Similarity Report is usually just that mechanism working: your references are formatted to the style guide, and other people cite the same work. What is worth looking at is where else the colour is. Matched text inside your own body paragraphs is a different conversation, and it is the one that needs an answer. Which matches count toward the number you were shown depends on the filters active on that report, which the report comparison covers.
If it was the AI Writing Report, your bibliography was not counted
A release note dated 9 August 2023 on Turnitin's AI writing detection model page states that bibliographies are now excluded when processing the AI writing report. That is not a filter somebody forgot to switch on. It is part of how the report is produced.
The model also only looks at what Turnitin calls qualifying text, meaning long-form prose made of standard grammatical sentences. Short non-sentence structures such as lists and bullet points fall outside the analysis, and a reference entry is not a sentence either way. Two consequences catch people out: the percentage is a share of the qualifying prose rather than of your whole document, and the amount of highlighting you can see will not reconcile with the number you were given. Both are unpacked in how to read a Turnitin AI writing report.
If your AI percentage came back higher than you expected, it came from your body text. The reference list is not where it came from.
Then why is the sentence with my citation in it highlighted?
Because the thing being scored is not the citation.
Turnitin describes the mechanism as cutting the submission into overlapping segments, giving each segment a score between 0 and 1, letting each sentence inherit the score of the segments it sits in through pooling, and aggregating those into a document-level number (AI writing detection FAQs). Nowhere in that description is there a step that recognises a citation, a quotation, or a source.
A sentence that happens to contain (Smith, 2019) is scored as prose, the same way the sentence before it was. What got flagged is the sentence. The citation inside it is incidental.
That rules out a whole family of attempted fixes. Changing the citation style, moving a parenthetical citation into a footnote, or rewording the attribution does not touch anything the model reacted to.
The most popular explanation for this is one Turnitin denies
In almost every thread on this subject, somebody explains that AI detectors measure burstiness and perplexity, and that academic prose scores badly on both because it is dense and evenly paced. It is a satisfying explanation. It is also not what Turnitin says it built.
The same FAQ addresses it directly: the model is transformer-based and is not explicitly programmed to evaluate specific signals such as burstiness or perplexity, and Turnitin adds that an individual prediction may not always be explainable in feature-by-feature terms.
That is worth knowing before acting on advice built on the other model. Deliberately varying your sentence lengths so your writing looks less even is a theory about a system whose vendor says it is not built that way. We go through where the burstiness story came from, and what replaced it, in what AI detectors measure beyond perplexity and burstiness.
Where Turnitin says its own false positives cluster
There is one more piece, and it is oddly specific. A release note dated 24 May 2023 on the same model page says Turnitin has observed a higher incidence of false positive detection in the first few or last few sentences of a document, and notes that these are often introduction or conclusion sentences written in a generic way.
Now think about where your conclusion sits. It is the last thing before your reference list.
If you saw colour near the bottom of a report and read it as your references being flagged, check where the highlighting actually stops. The paragraph above your bibliography is exactly the place Turnitin names as its own weak spot.
One related note if your document is short. Turnitin says predictions on documents of only a few hundred words tend toward all-or-nothing, because there is a single segment and no overlap to average across, so mixed content in a short document can come back scored as entirely AI.
What to check, in order
- Establish which report it is. If there is no number beside the AI indicator, only an asterisk, that is the AI report. Turnitin does not display a percentage for AI scores between 1 and 19, and says that band carries a higher false positive rate.
- If it is the Similarity Report, look at the matches outside the reference list. Those are the ones that mean something.
- If it is the AI Writing Report, read the highlighted prose and stop treating the percentage as a proportion of the document.
- Check whether the highlights sit in your first or last few sentences, the band Turnitin names as its own false positive hotspot.
- Remember what the vendor says the number is for. Turnitin states across several of its own pages that the indicator can be wrong and should not be the sole basis for action against a student. That is a reasonable sentence to bring to a conversation with a supervisor.
If the conversation has already started and you wrote the paper yourself, flagged, but you wrote it yourself covers what evidence is worth assembling.
If you decide to rewrite something
One warning before you paste anything into a rewriting tool. If the tool rewrites the whole file, your reference list sits inside its scope by default, and that kind of damage has been recorded.
Tadhg Blommerde, a lecturer at Northumbria University who reviews these tools with instructor-side Turnitin access and states that he takes no sponsorship and uses no affiliate links, built a test essay with in-text citations and an APA reference list and ran it through Walter Writes. In that February 2025 review he reports that the tool "even humanized the reference list", and separately points out a citation in the output that was not in the input (review video). An invented citation is a worse problem than the score you were trying to fix.
HumanPen is built around that constraint. Reference entries, in-text citation markers, footnotes, captions, figure numbering, table cells and cross-reference fields are protected and are not rewritten. Scope is something you set rather than something the tool assumes: select the passages yourself, or import the AI Writing Report and let its highlights define them, and the matched paragraphs are shown to you before anything runs. A DOCX or PPTX comes back as the same editable file. Nothing is made to score lower by inserting grammar or spelling errors. Eligible results can continue lowering AI for free.
None of that is a claim about what your next report will say. It is a description of what does and does not get touched.
Frequently asked questions
Does my bibliography raise my Turnitin AI score? No. Turnitin's release notes state that bibliographies are excluded when the AI writing report is processed, and the analysis covers qualifying prose in standard grammatical sentences, which reference entries are not. A high AI percentage came from your body text.
Should I delete my references before submitting so they do not get flagged? No. It does not change what the AI report analyses, since bibliographies are excluded from that processing anyway, and it hands in an incomplete paper to solve a problem you do not have. If you want to know what the report reacted to, look at which prose paragraphs are highlighted.
Why is my similarity score high but my AI score only an asterisk? They are independent measurements of different things. Turnitin shows an asterisk instead of a number for AI scores between 1 and 19, and attributes that to a higher false positive rate in that range. High similarity with an asterisk beside the AI indicator usually means a lot of matched text, often including the bibliography, and very little prose the model scored as AI.
Can I see the AI Writing Report myself? Not directly. Turnitin says the AI writing indicators are visible only to instructors and administrators. An instructor can download a PDF of the report and share it, which is worth asking for before you start reasoning from someone else's description of what got flagged.
KEEP READING