If I Rewrite One Flagged Sentence, Does the Score Drop by One Sentence's Worth?

Several accounts of how the Turnitin AI percentage is produced make it look additive. Each one can be traced back to a document, and in each case the document says less than the advice built on it.

HumanPen Team

· 22 min read

The short answer

No. The score a sentence carries was not earned by that sentence alone, so there is no exchange rate between sentences fixed and percentage points removed. Turnitin's published account is that sentences are extracted and grouped into overlapping segments, each segment is scored between 0 and 1, and a sentence inherits the score of every segment it sits in — which means a sentence near a boundary carries more than one score, pooled. That description names three operations — inherit, pool, aggregate — and publishes none of them, so there is no way to read one sentence's contribution out of it. Three specific accounts of that calculation are in circulation and all three are additive. Each one traces to a document, and in each case the document says less than the account built on it.

Those three are below, each next to its document. The mechanism itself, and the hand edits that follow from it, we wrote up separately in how to humanize AI text without a tool.

"The report is binary at the sentence level"

On a proofreading company's guide to reducing an AI score, retrieved 18 August 2026 and quoted exactly:

"There is no colour scale like the similarity report. The AI report is binary at the sentence level: each sentence is either flagged or not flagged. The overall percentage is simply the proportion of flagged text across the whole document."

As a description of the report in front of you, the first half is right. A sentence is highlighted or it is not; there is no gradient in the rendering. What it gets wrong is the thing behind the highlight. Turnitin's own account:

"Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score."

So the sentence did not earn a verdict. It inherited a probability, possibly two, from blocks of text that include its neighbours. The highlight is a rendering of a threshold, not the quantity the number is built from.

The last clause of that quote has a separate problem. "The proportion of flagged text across the whole document" is not what Turnitin says the percentage is a proportion of. The analysis covers only what it calls qualifying text — blocks written as standard grammatical sentences, which leaves lists and bullets out — and the consequence Turnitin draws is that "this percentage is not necessarily the percentage of the entire submission". If a third of your file is tables, bullets and a reference list, the denominator is not your document. We took that apart in how to read a Turnitin AI writing report.

What the advice built on it looks like: count the highlighted sentences, divide, work out how many you need to fix. Under the published description that arithmetic has no referent, because the highlight count and the percentage are not two views of one quantity.

"The average of all segment scores"

From a humanizer vendor's guide, which cites Turnitin's own help centre as its source for the mechanism. Retrieved 18 August 2026, quoted exactly:

"The model calculates an overall prediction based on the average of all segment scores, producing the percentage displayed in the AI writing indicator."

This one is closer, and the gap is narrow enough to be worth stating precisely. The published sentence is:

"These sentence scores are further aggregated and used to compute the overall document AI writing score."

Sentence scores, not segment scores. The segment score is an intermediate value a sentence inherits, sometimes twice; what gets aggregated is what comes out of the pooling.

Whether the two land on different numbers we cannot tell you, because Turnitin has published neither the pooling function nor the aggregation. What is different is what the operation ranges over, and that is the half a revision plan depends on. Under an average over segments, every segment counts the same however many sentences are in it, and each sentence sits in exactly one of them. Neither of those is what the published description says.

"Rewrite the first and last sentence of every paragraph"

The third account skips the mechanism and arrives as an instruction: rewrite the opening and closing sentence of every paragraph, because openings and closings get flagged most.

There is a real document behind it. In a release note dated 24 May 2023, under the heading "Improvements in detection at the start and end of documents", Turnitin wrote:

"Since launch, we have observed a higher incidence of false positive detection in the first few or last few sentences of a document. Many times these sentences consist of introduction or conclusion content written in a generic way. As a result, we have changed our detection logic to help reduce these false positives."

Three things to take from that, in order of how often they get dropped.

It says the start and end of a document, not of every paragraph. The advice has migrated down two levels of structure.

It is a fix announcement. The third sentence is the announcement. Citing the first two as a description of how the detector behaves now is citing a bug report as a spec.

And the note continues: "We also worked on making our segment boundaries detection more precise which could lead in some rare cases to change of boundaries compared with a previous version." The boundaries moved. Which sentence sits in which segment is a versioned implementation detail, not something to plan around.

If you are preparing for a misconduct conversation, do not take this release note in with you. The person opposite you can read the third sentence too. What you can reasonably say is narrower and holds up better: the vendor has revised the model in response to false positives, which is not the behaviour of something to be treated as settled evidence.

What the three have in common

Each one, if true, would let you compute what a single sentence is worth. That is what makes them attractive. A revision plan needs a unit of progress, and none of the published material supplies one.

You also cannot recover the missing number by experiment, for two documented reasons. Scores above 0% and below the 20% threshold are not displayed at all — an asterisk, no highlights — so a genuine move from 19% to 4% looks the same as no move. And in most setups the indicator is not yours to look at: Turnitin says the indicator and report are not visible to students, and that an instructor with the PDF download can share them with you.

One case does behave close to additively, and it is the one you least want. Where a submission runs to a few hundred words there is a single segment and no opportunity to overlap, so Turnitin's own account is that the prediction comes out mostly all or nothing — and that text mixing AI-generated and original content can therefore be reported as entirely AI-generated. That is not fine control. It is a switch.

What this changes about how you spend the evening

  1. Before you follow a rule about which sentences to rewrite, find the document it came from. All three above have one. In all three the document says less than the rule does, and the gap is visible in about five minutes of reading.
  2. Work from the highlights you actually have. If a report marked specific passages, those are the ones with evidence attached. The evidence behind a rule about position is a release note from 2023.
  3. Count the paragraphs you will have to re-verify, not the ones you will change. Every paragraph you touch is a paragraph you re-read for a broken citation, a shifted claim, or a number that moved. A plan that touches every paragraph in a thesis has quietly created a proofreading job the size of the thesis.
  4. Keep the file you started with. The original and the intermediate drafts are the part of this with evidential value, and they keep it whatever the number does. The mechanics are in version history as evidence.

The tool we build, and the number it cannot give you

HumanPen takes the AI Writing Report as a PDF and works from the passages it marks. One of its rules is relevant here: a paragraph is the smallest unit it will rewrite, so a selection covering part of one is widened to the whole paragraph before the job runs.

That rule is not a claim about the scoring. It is there because a highlight is free to start mid-clause, and rewriting exactly the highlighted characters hands you a paragraph written by two people — half in one register, with the second half leaning on terms introduced in the half nobody touched. Turnitin has not published what unit a rewrite would need to cover to move a score, and we are not in a position to work it out from the outside.

What we cannot give you is the exchange rate, and neither can anything else. A tool that offers you a predicted percentage is offering a number it does not have.

Frequently asked questions

If a sentence is highlighted, does removing it remove its share of the score? Not predictably. The highlight reflects a score the sentence inherited from the segments around it, and removing text changes the segments. The document number is an aggregate of pooled sentence scores, not a sum of highlighted sentences.

Does the score fall by the same amount if I rewrite a whole paragraph? There is no published amount, for a paragraph any more than for a sentence. What is published is that segments overlap, so what happens to any one passage depends on the text on either side of it.

Why do the highlights and the percentage disagree? Because the percentage is calculated over qualifying prose only, and Turnitin says a document mixing several writing types "would result in a disparity between the percentage and the highlights."

Is the first and last sentence of each paragraph really flagged more? The document that started this said the first few and last few sentences of a document, said it in 2023, and in the same note announced a change to the detection logic to reduce it. It is not a current description of behaviour and it is not usable as a defence.

Can I test my edits myself? Usually not. Anything under 20% shows as an asterisk with no number and no highlights, and in most configurations the indicator is visible to instructors and administrators rather than to you.

---

KEEP READING