The integrity flags on a Similarity Report: hidden text and replaced characters
There is a second tab on the report that has nothing to do with how much your text matches. It is worth understanding before you let any tool near a manuscript you are about to submit.
HumanPen Team
· 7 min read
The short answer
The Similarity Report has a Flags tab separate from the percentage. It marks two documented patterns: hidden text — white-on-white characters, concealed quotation marks — and replaced characters, where letters are swapped for lookalikes from other alphabets. The documentation states that lookalike characters are automatically swapped back before scanning, so the similarity match happens anyway. A flag is explicitly not a finding of wrongdoing, but it does put a red marker on your submission for a human to look at.
What the Flags tab is
The guidance describes it in one sentence: "Turnitin's algorithms look deeply at a document for any inconsistencies that would set it apart from a normal submission. If we notice something strange, we Flag it for you to review."
If anything is detected, a number appears next to the Flags tab above the submission — 1 when one flag type is present, 2 when both are. Opening the tab gives an Integrity Flags for Review panel, and in the document itself "a red flag with a number inside indicating the flag type and red boxes around all suspicious text related to that flag."
The caveat sits right there in the source, and it matters in both directions: "A flag is not necessarily an indicator of a problem. However, we recommend you focus your attention there for further review."
So it is not an accusation. It is a red box drawn around part of your manuscript with an instruction to the reviewer to look harder at exactly that part.
Replaced characters, and why the trick fails twice
This is the one worth reading carefully, because it is the technique most likely to be inside a tool rather than typed by a person.
"Some characters in different alphabets can look similar enough that to the naked eye, it is difficult, if not impossible, to tell them apart." A Cyrillic character that renders identically to a Latin one, dropped into the middle of a word, breaks the string a matching engine is looking for while leaving the page looking normal.
The documented response has two parts, and the first is the part nobody mentions:
"Turnitin automatically swaps these characters out when scanning a submission so they will not affect the Similarity Report. However, by replacing characters, the intent is to try and interrupt a similarity match."
So the substitution is reversed before comparison — the match still happens — and the attempt is raised as a flag. The technique does not reduce the number it was aimed at, and it adds a marker that was not there before.
That is a rare thing to be able to say with a citation rather than an opinion: a documented bypass technique that is documented not to work.
Why this is a question about tools, not about people
Most authors are not going to sit down and paste a Cyrillic 'о' into every third word. But a piece of software will, invisibly, in a second, across a whole manuscript — and you will not see it, because the entire point of the technique is that it renders identically.
Which makes this a question to ask before you hand a file to anything: does what comes back contain characters that were not in what went in?
You can check without any special tooling. First, though, rule out the version of the check that sounds obvious and does not work.
Counting how many times a common letter appears tells you nothing. Rewriting tends to lengthen text, and the letters go up with it. On our own before-and-after files the plain Latin "a" went from 4,631 to 5,092 and the "e" from 6,970 to 8,005, with no character substituted anywhere. Count letters, find more of them, and all you have learned is that the document got longer.
Compare inventories instead of counts. Export the plain text of the file you sent and the plain text of the file you got back, list the distinct characters that occur in each, and look for anything in the second list that is missing from the first. Cyrillic and Greek letters that look Latin. Zero-width spaces. Soft hyphens. Counts changing is expected. A character appearing that was never in your file is not.
And two structural points about who carries the consequence. The flag lands on the submission, which has your name on it. And the tool that introduced the characters is not in the room when the report is read.
What we do and do not do with your characters
We are in the category where this technique lives, so the honest thing is to be explicit about it rather than let it go unsaid.
HumanPen rewrites sentences. It does not insert invisible characters, it does not swap letters for lookalikes from other alphabets, and it does not colour text white. There would be no point: the documentation above says the substitution is reversed before matching and flagged on the way past, so the only thing such a technique reliably produces is a red box on your manuscript.
We ran the inventory comparison on our own output before publishing this, on four before-and-after pairs. The set of distinct characters in each output was a subset of the set in the input every time. Nothing new, no Cyrillic, no Greek, no zero-width anything. The count of non-ASCII characters went down rather than up: em dashes 58 to 13 in one file, curly apostrophes 41 to 6. We also checked that the comparison can see what it is looking for, by planting one Cyrillic "о" and one zero-width space into a copy of the output and running it again, which listed both.
Those are our files and we do not publish the corpus, so take the numbers as ours and the procedure as yours. Run it on your own pair. Including on anything we hand back.
Frequently asked questions
Does a flag mean I have been accused of something? No. The documentation states that a flag "is not necessarily an indicator of a problem", while also recommending that a reviewer focus attention there. It is a prompt to look, not a finding.
Do lookalike characters lower a similarity score? Not according to the documentation, which says those characters are automatically swapped out when a submission is scanned so they will not affect the Similarity Report. The attempt is also raised as a flag.
How many flag types are there? Two are documented — hidden text and replaced characters — and the tab shows a 1 or a 2 depending on how many types were detected.
Can I check my own file for this before submitting? Not through the report, which is generated on the reviewer's side. You can inspect the file itself. List the distinct characters in the version you sent out and in the version that came back, and look for any character in the second list that is absent from the first. Do not count letters instead: rewriting makes text longer, so letter counts rise on their own. Also check for text set in white or in a colour matching the background.
KEEP READING
Why your quoted and cited text still shows up in the Similarity Report
9 min read
Academic writingWho runs the similarity check on your manuscript, and what it compares against
8 min read
Academic writingFinal manuscript check: AI disclosure, word count, citations and DOCX integrity
12 min read