Still flagged after a Turnitin recheck? Compare two reports, not just two scores

A second percentage cannot tell you whether the first revision failed. The document, model and denominator may all have changed. A useful comparison fixes those variables, maps highlights by passage and separates persistent cores from moved boundaries, new flags and report-version effects.

HumanPen Team

· 8 min read

The first principle: compare locations before totals

Suppose a first Turnitin AI Writing Report shows 42% and a recheck shows 28%. It is tempting to call that a 14-point improvement and stop. The number alone cannot support that conclusion. It does not tell you whether old highlights disappeared, their boundaries moved, a different section became highlighted, or the denominator changed because the document structure changed.

The reverse is also true. A percentage can stay flat while the report points to completely different passages. It can rise even though several old highlights vanished, if the revised document contains less qualifying prose or a model update classifies another region differently. A total compresses all of those events into one value.

The useful question is therefore not "Did the score fall?" but "What happened to each previously highlighted passage, and under what report conditions?" Answering it requires the two source documents and two full PDFs. Screenshots of the cover percentages are insufficient.

A recheck is a new observation, not a controlled experiment by default. You must establish what changed in the document and what changed in the detector before attributing the result to revision.

The comparison method below is HumanPen's operational workflow, derived from Turnitin's documented report behaviour. It is not an official Turnitin procedure and does not claim that passage matching can reveal why a model changed its classification. Its purpose is narrower: stop you from rewriting the whole paper when the evidence points to a small, identifiable remainder.

First decide whether the reports are comparable

Before drawing arrows between highlights, complete a comparability check. If one of the rows below cannot be confirmed, record the uncertainty instead of treating the two percentages as an exact before/after measure.

VariableWhat to verifyWhy it can change the result
Document identityKeep the exact first and second DOCX/PDF files; record filenames, saved dates and word counts.Comparing a different export, appendix set or draft confuses revision effects with file differences.
Submission languageConfirm English, Spanish or Japanese, and note mixed-language passages.Language models and capabilities differ; paraphrasing and bypasser detection are English-only.
Report generation dateRecord when each report was generated and any update notice.Turnitin model updates are not retroactive; resubmission is required for updated scores.
Report stateDistinguish a number, 0%, `*%`, processing error, ineligible file or updated-report notice.An asterisk hides the above-0-and-below-20 range and supplies no highlights, so it is not a numeric zero.
Qualifying proseNote major changes to tables, bullets, code, references or other non-prose content.The AI percentage uses qualifying long-form prose, not the visible whole-document word count.
Category legendRecord which categories each report actually shows, and the exact labels on them.The categories themselves changed. On 4 August 2026 Turnitin merged the purple AI-paraphrased category into a single blue AI-generated one, so a report from before that date and a report from after it are not showing the same legend.

Turnitin's model release notes repeatedly state that model updates are not applied retroactively and that a resubmission is needed for an updated score. Updates on 14 October 2025 and 12 February 2026 aimed to improve recall, while the May 2026 update affected the Spanish model. Two reports generated on opposite sides of an update do not isolate the effect of your edits.

The legend has its own history. Turnitin changed segment-boundary logic in May 2023, added paraphrase-related processing without a visual change in December 2023, redesigned categories and low-score treatment in July 2024, and on 4 August 2026 combined the two AI categories into one, saying that the purple highlights previously used for AI-paraphrased text will no longer be shown. Turnitin adds that only files submitted after that date get the new layout; anything already in the system keeps the report it has until it goes through again. Boundary movement and a changed legend can therefore arise from the report version, not only from a changed sentence.

Classify every highlight into one of six outcomes

Once the reports are comparable enough to proceed, stop thinking in pages. Page numbers move when paragraphs reflow. Use a stable passage anchor instead: section heading plus the first eight to twelve words of the original passage. Then assign each highlighted region an outcome.

OutcomeWhat it looks likeReasonable next action
Persistent coreSubstantially the same sentences remain highlighted in both reports.Review meaning, drafting history and the actual revision. If another revision is permitted, limit it to that stable core.
Boundary movedThe same topic remains flagged, but highlighting begins or ends one or two sentences earlier or later.Read the entire paragraph. Do not treat every newly coloured boundary sentence as a separate failure.
Category switchedThe passage stays highlighted but its category changes, or the second report shows one category where the first showed two.Record the switch and both report dates. Since 4 August 2026 the report shows a single merged category, so a switch across that date is a legend change. Do not infer a new tool was used between submissions.
New highlightA passage unhighlighted in report one is highlighted in report two.Check whether the passage itself changed, moved next to revised prose, or was processed by a different model. Review it independently.
Highlight disappearedA previously highlighted passage has no corresponding highlight in the second report.Mark it resolved for this comparison. Do not keep rewriting it simply because the total remains above a preferred number.
Not comparableThe passage was deleted, replaced wholesale, became non-qualifying, or one report provides only `*%` without highlights.Document why no location-level comparison is possible; do not classify it as success or failure.

A moved boundary is especially easy to overreact to. Classification models operate on textual context, so editing one sentence can change how neighbouring sentences are grouped. Turnitin's history also includes explicit changes to segment-boundary precision. The correct unit for review is the coherent paragraph, not each coloured line treated as independent evidence.

A new highlight is not automatically newly written AI text. It means the second report surfaced a region the first did not. Document edits, context, qualifying-text changes and model version are all competing explanations.

Build a comparison sheet that another person can audit

A useful comparison sheet is small enough to understand and precise enough to reproduce. Create one row per passage anchor, not one row per highlighted line. Assign IDs such as `P01`, `P02` and keep the first report as the coordinate system.

FieldExamplePurpose
Passage ID and anchorP03 · Methods · `Participants completed the survey...`Finds the same prose even when pagination changes.
Report one stateAI-generated · paragraph 2 to 4Preserves the original category and extent.
Revision madeRebuilt claim order; retained sample size and citationSeparates substantive editing from synonym replacement.
Report two stateAI-generated · paragraph 3 onlyRecords the current evidence without erasing the first result.
OutcomeBoundary moved; persistent coreProvides a consistent comparison label.
DecisionReview paragraph 3; freeze paragraphs 2 and 4Turns the report into a bounded action rather than a whole-document rewrite.

Use Word's document comparison or tracked changes to establish what text actually changed between source files. Do not rely on memory. For each row, verify numbers, citations, quotation status and the strength of the claim before considering style. A lower AI classification is not an improvement if a causal claim became correlationally wrong, a limitation disappeared or a citation no longer supports the sentence.

Keep the comparison sheet private with the source documents unless the applicable process requires disclosure. It may contain unpublished research or student information. If it must be shared, remove unrelated personal data and retain only the evidence needed to explain the passage history.

The sheet is also useful when no further editing is appropriate. It can show that report two flagged different material, that a model update separates the runs, or that the new result is an asterisk with no location evidence. A well-documented "cannot compare" is more accurate than a confident but invented improvement percentage.

Why the total can move in unintuitive ways

The AI Writing Report guide says the percentage covers qualifying prose and can visibly disagree with the amount of highlighting when a document mixes writing types. Editing the document can change both the numerator - classified qualifying text - and the denominator - all qualifying text. That makes simple percentage-point arithmetic incomplete.

  • Same total, better-localised result: three old highlights disappear and one equally sized new region appears. The total hides a complete redistribution.
  • Lower total, unresolved core: peripheral highlights disappear but the central disputed passage remains. The number falls while the most important review question survives.
  • Higher total, fewer visible pages: non-prose or other qualifying-text changes shrink the denominator. A larger percentage need not cover more of the physical document.
  • *Numeric result to `%`:** the second result is somewhere above zero and below 20, but Turnitin does not surface the exact score or highlights. It is not valid to record it as zero or calculate the exact reduction.
  • Category redistribution: in a report issued before 4 August 2026, the two category shares move while the combined total barely changes. Since that date there is one category, so this reading only applies to older reports, and a pair straddling the date cannot be compared at category level at all.

This is also why there is no honest formula for "percentage of the revision that worked". The detector does not track edits or assign causal credit. You can report observable passage outcomes and total changes, but you cannot turn them into a verified effectiveness rate for the editing process.

Turnitin itself recommends treating the result as one data point alongside knowledge of the writer, the work and institutional policy. Its review guidance directs attention to the highlighted categories and human judgment rather than a score in isolation.

Know when to stop revising, and preserve the record

Repeated rewriting has costs. Meaning drifts, citations detach from claims, discipline-specific terms become vague, and the author's voice can disappear. A report gives no instruction to chase zero, and `*%` explicitly withholds an exact low-range result because Turnitin considers that range less reliable.

Stop and review rather than automatically revise again when any of these conditions applies:

  • A formal review is open. Preserve originals and follow the institution's process before modifying a copy.
  • The reports are not comparable. Different model eras, file versions or report states make another edit unable to answer the uncertainty.
  • Only boundaries moved. The remaining difference may reflect contextual grouping rather than a stable new passage to rewrite.
  • Further changes would weaken accuracy. Freeze numbers, quotations, methods, definitions and claims whose exact wording carries substantive meaning.
  • The current prose is owned and defensible. You can explain it, support it from drafts and sources, and it meets the applicable AI-use and disclosure rules.
  • The only target is a guaranteed score. No model, service or revision method can promise the output of a future detector version.

Keep a compact evidence package: both source documents, both full reports, the comparison sheet, a change log, and the policy or assignment instructions that governed the work. Name files with dates and never overwrite the originals. If a third check occurs, add it as a new observation instead of replacing report two.

If revision is permitted and report two contains a small set of persistent cores, the current document and current report are the inputs that matter. HumanPen intentionally works from the latest highlights rather than carrying forward everything report one once flagged. That scope control prevents resolved passages from being rewritten again, but it remains a revision aid, not a promise that report three will be blank.

The defensible outcome is not "the number went down". It is "these passages disappeared, these remained, these moved, these were new, and these variables prevent a clean comparison".

KEEP READING