I Edited My Paper and the AI Score Went Up
The number is not built by adding up sentence-level verdicts, so nothing in how it is produced makes it fall in step with the sentences you fixed. Three published parts of the calculation can push it the other way while your edits do exactly what you meant them to.
HumanPen Team
· 20 min read
The short answer
A second report that reads higher than the first is not a measurement of how well you revised. Three published parts of how Turnitin produces the percentage can each move it upward on their own. Sentences are scored inside overlapping segments, so an edit changes what its neighbours are grouped with, and the effect has no fixed radius. The percentage is calculated over qualifying prose only, so a revision that takes prose out of the document shrinks the denominator while the flagged text sits still. And the highlights you can count and the number on the cover are different quantities, which are allowed to move in opposite directions in the same report.
None of that says the revision was wrong or that the writing got worse. It says the reported quantity does not behave the way the word "percentage" invites you to assume. Below is each mechanism next to the document it comes from, plus a fourth that only bites if the revision also made the document much shorter, and then what a rise does and does not establish.
Nothing you edit stays local
Start with how a sentence gets a score, because the answer is that it does not get one of its own. Turnitin's FAQ describes the pipeline in one paragraph:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."
Two words there carry the whole problem: "overlapping", and "pooled". A sentence sits inside more than one segment, and each segment is classified as a block. So every score a sentence carries was produced partly by the sentences on either side of it.
Rewrite one sentence and you have changed an input to the scores of its neighbours, whether or not you touched them. A paragraph you deliberately left alone can come back highlighted because the segment containing it now contains different text. The reverse happens too, which is the version nobody writes in to ask about.
The boundaries themselves are a moving part in the vendor's own account. In a release note dated 24 May 2023, describing work done to reduce false positives at the beginnings and ends of documents, Turnitin added:
"We also worked on making our segment boundaries detection more precise which could lead in some rare cases to change of boundaries compared with a previous version."
That was a 2023 improvement rather than a description of a current fault, and it is worth reading for one narrow reason: segment boundaries are part of the process, and Turnitin has documented changing them. The FAQ also makes the existence of pooling and aggregation explicit. What it does not publish is the pooling rule, any weighting, or the aggregation algorithm. So there is no way to work out beforehand which neighbours a given edit will move, or in which direction.
The denominator can move while you are looking somewhere else
The second mechanism has nothing to do with classification at all. It is arithmetic, and it is the one that produces the largest surprises.
"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures. This percentage is not necessarily the percentage of the entire submission."
Put numbers on that. Say the first version held 6,000 words of qualifying prose, of which 1,800 came back flagged: 30%. For the second version you rewrite 300 flagged words and they stop being flagged, leaving 1,500. In the same sitting, because a supervisor said the discussion was heavy going, you turn 2,400 words of it into a bulleted summary and a table of figures. Qualifying prose is now 3,600 words. 1,500 over 3,600 is a shade under 42%. The score climbed twelve points in a revision where less text was flagged than before and nothing new was flagged anywhere.
| A change you may have made | What it does to the qualifying prose |
|---|---|
| Turned three paragraphs into a bulleted list | Those words leave the analysis. Turnitin names bullet points and other short non-sentence structures as things it does not analyse |
| Moved a description into a table | Depends what went in. An August 2023 release note says long-form prose in tables is now processed; a table of numbers is not prose either way |
| Added twenty more references | Nothing at all. The same release note says bibliographies are excluded when the AI writing report is processed |
| Cut 1,500 words of unflagged prose to meet a limit | The denominator falls and the flagged share rises, with no change whatsoever to the flagged text |
Both of those August 2023 lines close the same way, telling you to resubmit an older paper before either change applies to it, so what they describe is how a report gets produced now rather than how an old one was. And none of the table is advice. Two of those rows describe making a document harder to read, and a marker who receives a discussion section as bullet points notices. The reason to know the list is diagnostic: if one of those rows describes your revision, the movement has an explanation that says nothing about how your sentences read. The reference-list case has its own tangle, which we took apart in why Turnitin flagged my references.
Fewer highlights and a lower number are not the same event
Most people check the highlights first, because a shorter list of them is the part that feels like progress. Turnitin says plainly that the two can disagree:
"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."
So "I have fewer highlighted passages than last time and the number still went up" is not a contradiction in need of resolving. It is two quantities with different denominators being read as one. And under the 20% threshold there is nothing to count: Turnitin withholds both the score and the highlights for results in the 1% to 19% range, so a report with no highlights on it is not a report with nothing flagged. We went through the parts of the report one at a time in how to read a Turnitin AI writing report.
A shorter submission is a less stable one
If the revision was also a cut, there is a fourth mechanism, and it applies to short submissions in particular:
"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap. This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."
The overlap that smooths a long document out is not available on a short one, and the failure mode Turnitin describes runs upward: mixed content flagged as entirely AI. There is a floor underneath that as well, since the file requirements ask for at least 300 words of prose before a report is generated at all, which we covered in can Turnitin check a whole thesis. Between the two, a revision that shortened the piece and a revision that changed its prose are two experiments run at once, and the report hands back a single number for both.
What an upward move does and does not establish
Nothing in the report records what you changed. There is no field attributing the movement to a particular edit. So the rise is an observation about two files, not a judgement on the work between them.
- It does not establish that your edits made the writing look more like AI. Under the mechanisms above, the number can rise while every sentence you rewrote comes back clean.
- It does not establish that the first number was the trustworthy one. The two reports are separate outputs, and neither explains why it differs from the other.
- It does not establish that another rewriting pass is the next move. If the denominator moved, a further pass over the prose changes only the numerator.
There is also a sentence on Turnitin's FAQ that belongs in this conversation, and it is usually quoted with its second half missing:
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas. If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
Read to the end and it is not an accusation pointed at you. It is the vendor naming kinds of human writing its own model can flag incorrectly, then telling whoever reads the report to take that possibility into consideration when looking at the percentage. If a supervisor is looking at a number that went up, that caution sits on the same page as the number. The hand-editing list built on the first half is in how to humanize AI text without a tool, and the method for comparing two reports passage by passage instead of total against total is in still flagged after a Turnitin recheck.
Where a report-scoped rewrite fits
All of the above is an argument about scope. If changing a sentence changes what its neighbours are grouped with, then the smallest sensible thing to change is larger than a sentence, and the passages you are happy with are worth leaving exactly as they stand.
That is the scope HumanPen describes on its public report page. Upload the DOCX with a Turnitin or iThenticate AI report, and the flagged passages define what gets rewritten. Everything else stays frozen. The public humanize page also states that billing is based on the words actually rewritten, not the size of the file you sent.
If a later report still flags passages, eligible results can continue lowering AI for free. That service rule is not a forecast. It does not say what the next report will read, and Turnitin has published no mapping between an edit and a movement in the score.
Frequently asked questions
Can a Turnitin AI score really go up after I revise? Yes, and there is nothing unusual about it. The percentage is not assembled from per-sentence verdicts, so it has no reason to track the number of sentences you fixed. Overlapping segments, a changed amount of qualifying prose, and a shorter document each move it on their own.
Does a higher number mean my edits made things worse? Not by itself. The report contains no record of what you changed and no attribution of the movement to any edit. A rise is consistent with edits that worked, edits that did nothing, and a document whose structure changed underneath the calculation.
I have fewer highlighted sentences but a higher percentage. Which one is right? Both, because they are not two views of the same quantity. Turnitin says a document mixing writing types will show a disparity between the percentage and the highlights, since the percentage is computed over qualifying prose while the highlighting sits on the whole file.
I deleted a lot of text to hit a word limit and the score jumped. Why? Most likely the denominator. The percentage runs over qualifying prose, so removing unflagged prose raises the flagged share without any change to the flagged text. If the result is now only a few hundred words, Turnitin also says short documents are predicted mostly all or nothing.
Should I just rewrite the whole paper? That is the response the number invites and it is rarely the one the evidence supports. Compare where the highlights are in both reports before deciding how much to touch, and freeze anything whose exact wording carries meaning, such as numbers, quotations, definitions and method steps.
KEEP READING