Why your quoted and cited text still shows up in the Similarity Report

A properly quoted, properly cited block appearing as a match feels like a bug. It usually is not. There is a documented recognition step in between, and it is the part that can miss.

HumanPen Team

· 9 min read

The short answer

Every match gets sorted into one of four groups according to whether the software detected quotation marks and an in-text citation around it: not cited or quoted, missing quotations, missing citation, or cited and quoted. That detection is a machine learning model with published limits — it is trained on English, its citation recognition covers four named formats, it recognises seven specific quotation-mark shapes, and it can miss a citation placed far from the matched text. A correctly attributed passage in the wrong shape lands in the wrong group.

The percentage first, because it is not what people assume

The overall similarity figure has a definition you can hold in your head: matching words divided by total words in the document. The formula weights nothing and discounts nothing for being attributed.

What moves the number is a different mechanism entirely, and it is not yours to set. The same help site has a guide to exclusion filters, and it lists five: exclude bibliography, exclude quoted text, exclude cited text, exclude small matches, exclude small sources. The one that matters most here is the second. It "ignores text enclosed in quotation marks or formatted as block quotes, preventing properly quoted material from appearing as matches", and it also covers indented blocks when the file is a .doc or .docx.

Which is already half an answer to the question in the title. Sometimes a correctly quoted paragraph shows up because nobody turned that filter on. Whether it was on is a decision made by whoever ran the check, and it does not appear on the number they send you.

The denominator is worth pinning down, because the AI writing indicator on the same platform does not use that one. Its percentage is calculated over qualifying prose rather than over your whole file, which is why that number and its highlights can look out of proportion to each other. Two numbers, two denominators. We set the two side by side in reading a Similarity score next to an AI writing score.

One threshold worth knowing before anything else: a submission needs at least 20 words to produce a Similarity Report at all.

And the framing sentence the documentation opens with, which is easy to skip past: "Highlighted matches are instances of text similarity; they do not always indicate plagiarism. The match could be a quote or cited material listed in a bibliography."

The four groups

Match Groups split the overall similarity by how the matched text was attributed. In the documentation's own wording:

  • Not Cited or Quoted — matches "that are not written as a quotation or has no citation to its original source".
  • Missing Quotations — "This text is cited, but the match is so exact that it may also require quotation marks."
  • Missing Citation — "This text is written as a quote, but lacks a citation to its original source."
  • Cited and Quoted — "This text contains a quotation and is cited to a source."

The stated purpose is to let a reviewer "distinguish between integrity issues, teachable moments, and properly-included source text". Which means the group a match lands in is doing real work in how a person reads your manuscript — more work, arguably, than the headline percentage.

So the practical question stops being "why did this match" and becomes "which group did this match land in, and is that the right group".

How it decides, and where it misses

This is the part that is documented and almost never repeated anywhere else.

Citation and quotation recognition is a trained model, not a rule: the guidance says machine learning technology is used "to recognize in-text citations and quotation marks associated with matched text". And the vendor states its own error rate in plain language: "Since there are innumerable ways to include in-text citations and quotation marks, the Similarity Report won't get it right every time."

Three documented limits follow from that.

It is English-only. "The technology that powers Match Groups is trained on English content, and this functionality is currently only available for submissions in English."

Its citation training covers four formats. "Citation recognition models have been trained on citations in certain formats: APA, MLA, Turabian, and IEEE numbered citations and references. If a citation is included for a match but is not in one of these formats, the report may--but is less likely to--recognize those as citations."

Read that list against your own discipline. Vancouver, Chicago author-date, Harvard and its many house variants are not on it. That does not mean they are never recognised — the wording is "less likely" — but if you write in a style that is not one of those four, the odds that your attributed quotation is read as attributed are documented to be lower. Which style you are actually being held to is its own mess, and we went through it in which Harvard referencing style is yours and choosing a citation style by discipline.

Distance matters. "Citations may not be recognized if they are placed very far from the matched text." A block quote whose citation sits at the end of the following paragraph is exactly this case.

The quotation marks it recognises

Seven shapes, listed explicitly on the Match Groups page:

  • Straight double quotes: "…"
  • Curly single quotes: ‘…’
  • Guillemets: «…»
  • Reversed guillemets: »…«
  • Low-high double quotes: „…“
  • White corner brackets: 『…』
  • Corner brackets: 「…」

Copy that list carefully, because the entry for double quotes is the straight one and the entry for single quotes is the curly pair. And the exclusion-filters guide on the same help site prints a different list again: eight entries, with straight single quotes rather than curly, 《…》 and 〈…〉 added, and 「…」 gone. Two published lists on one help site, ten distinct shapes between them, and neither page mentions the other.

Do not turn that into a rule about which shape to use. Turn it into a rule about not letting anything change your quotation marks between the draft you proofread and the file you upload.

And one failure mode is stated outright: "Quotation marks are not recognized if they are separated from the text by a space (for example, " this improperly quoted text")."

That is the single most checkable thing in this article. A stray space after an opening quotation mark — the kind that survives a copy-paste out of a PDF, or gets introduced when a paragraph is reflowed — is enough to move a match from Cited and Quoted into Missing Quotations.

Where quotation marks are not recognised, the match falls to "Missing Quotations" or "Not Cited or Quoted", depending on whether a citation was detected. The two failures compound.

Which source gets named when several match

When the same words match more than one source, the report picks one to show first, and the priority order is published: Internet, then Publications, then Submitted Works.

That ordering can be counter-intuitive when you are looking at your own work. A sentence you took from a published paper, correctly, may be attributed in the card to a website that reproduced that paper. There is a "Show overlapping sources" control that reveals the rest, and a View other sources action on an individual card.

Worth knowing before you argue with a match: the source named on the card is the strongest match, not necessarily the source you used.

A twenty-minute pass over your own manuscript

None of this requires access to the report. You can do it in the file you are about to submit.

  1. Find every direct quotation and check the opening mark sits flush against the first character, with no space between them. Search for a quotation mark followed by a space.
  2. Check the mark shapes, and check them again after any conversion. A round trip through a template, a PDF or a co-author's editor swaps straight for curly without telling you. In a translated manuscript you may have shapes that are on neither of the two published lists.
  3. Move citations closer. If a quotation's citation lives at the end of the next sentence or paragraph, bring it adjacent to the quoted material.
  4. Note your citation style. If it is not APA, MLA, Turabian or IEEE numbered, expect recognition to be less reliable and be ready to explain a match rather than assume the report will classify it for you.
  5. For anything still matched, check the source card rather than the highlight colour, since the named source is the strongest match and may not be your actual source.

The line between the two halves of a report

The similarity side of a report is not our department, and rewriting is not what fixes it. If a passage matched because it is a quotation, the answer is attribution. Changing the words of a quotation just makes it a worse quotation.

HumanPen works on the other indicator. Where a manuscript comes back with an AI writing percentage and specific passages highlighted, importing that report is what tells it which paragraphs to open.

If your report has both numbers on it, they are separate problems with separate fixes, and treating them as one is how people end up rewriting quotations that only needed a pair of quotation marks moved.

Frequently asked questions

My quotation is cited properly and it still shows as a match. Is that wrong? Not necessarily. The documentation states that highlighted matches are instances of text similarity and do not always indicate plagiarism, and there is a separate exclusion filter for quoted text that somebody has to switch on. With it off, correctly quoted material is expected to match. What matters then is which of the four match groups it lands in.

Does the report understand my citation style? Citation recognition is documented as trained on APA, MLA, Turabian and IEEE numbered formats. Other formats may still be recognised, but the guidance says it is less likely.

Why did my quotation marks not count? The recognised shapes are listed, and there is one explicit failure: marks separated from the text by a space are not recognised. Check for a space immediately after an opening quotation mark.

Does any of this apply to a manuscript not written in English? Match Groups are documented as trained on English content and available only for English submissions.

KEEP READING