Who runs the similarity check on your manuscript, and what it compares against

Authors tend to picture the check as a search engine looking for their sentences on the open web. The arrangement behind it is more specific than that, and a few of its details change what a matched passage actually means.

HumanPen Team

· 8 min read

The short answer

If your journal belongs to Crossref and uses Similarity Check, the check is a service Crossref provides to its members at reduced rates, running on iThenticate from Turnitin. Whether your manuscript goes through this particular check depends on whether the publisher participates, so that is the fact to establish first. Where it does apply: the editor uploads the manuscript and gets a report, you do not run it and normally do not see it unless it is sent to you, and Crossref says participating members get read-only access to the full text of articles in the Similarity Check database for comparison, describing that database as containing over 78 million full-text scholarly content items.

Who actually runs it

Crossref describes the service in a single line: "A service provided by Crossref and powered by iThenticate—Similarity Check provides editors with a user-friendly tool to help detect plagiarism."

Two words in that sentence carry most of the meaning for an author.

Editors. The tool is described as being for editors, who "upload a paper" and get back "a report highlighting potential matches and indicating if and how the paper overlaps with other work". The stated purpose is to "assess the originality of the work before they publish it". You are not a user of this software; you are its subject.

Members. Similarity Check is a Crossref member service. Crossref says the arrangement gives members "reduced-rate access to the iThenticate text comparison software from Turnitin", and that "Only Similarity Check members benefit from this tailored iThenticate experience". So whether your manuscript passes through this particular check at all depends on whether the publisher is participating, which is a fact about the publisher rather than about your field.

That also means "the journal ran a plagiarism check" describes at least two different situations: a publisher inside this scheme, and a publisher using something else entirely. If you are told a number, it is fair to ask which one produced it.

One thing to keep in mind while reading any of the above: Crossref is the organisation providing this service to its members, with discounted checking fees attached. Its descriptions are accurate about what the arrangement is, and they are not an independent assessment of whether the check is worth running. Every quotation in this article is a description of the mechanism, not an evaluation of it.

What it compares against

The comparison set is the part authors most often guess wrong.

Crossref describes the database as containing "over 78 million full-text scholarly content items", and describes what membership buys as including "read-only access to the full text of articles in the Similarity Check database for comparison purposes".

Full text, not abstracts and not metadata. The iThenticate documentation describes the search targets more broadly still — depending on the settings chosen, they may include "billions of pages of active (and archived) internet information, previously submitted works" and "tens of thousands of periodicals, journals, and publications".

Two consequences worth holding onto:

  • A paywalled paper is inside the comparison set. Being unable to read something yourself says nothing about whether your sentences will be compared against it.
  • The set is chosen per check. Search targets are described as selected for the assignment, so two editors running the same manuscript can be running it against different targets. A number from one is not a prediction of a number from the other.

What the report is for, and what it is not

The iThenticate guidance opens by warning that a highlight is a statement about text overlap and nothing more — a properly quoted passage or an entry in your bibliography is expected to light up.

The percentage itself is arithmetic, and the documentation says so: count the words that matched something, then divide by how many words the manuscript has. The formula discounts nothing for being attributed.

Which is not the same as saying your reference list is bound to inflate it. What decides that is a filter, and iThenticate's guide to them describes an exclude-bibliography setting that "automatically identifies and ignores reference lists or bibliographic sections so they don't inflate the similarity percentage", plus separate switches for quoted text and for cited text. So the same manuscript can produce a percentage with the bibliography in it or a percentage with the bibliography out of it, and the number itself does not tell you which one you are holding. That is the first thing to ask about any figure somebody quotes at you.

That is why the report sorts matches into groups by whether it detected quotation marks and an in-text citation around them — a recognition step with its own documented limits, and the reason two manuscripts with the same percentage can mean completely different things.

Where your own published work sits

The database is built from members' registered content. Crossref does not spell out the consequence for authors on either of its Similarity Check pages, so treat the next step as inference rather than quotation: if you have published with a participating member, your own articles should be expected to sit in the corpus that later manuscripts are compared against, including your own next one.

That is the ordinary explanation for a puzzle authors hit the second time round — a chunk of your new paper matching a paper you wrote. It is not a defect in the check. It is the check working on a corpus that has you in it.

Nothing about that decides whether reusing your own text is acceptable in your field. It only means the report will show it, and that you should expect to explain it rather than be surprised by it.

The version question, briefly

Crossref's documentation lists setup paths for iThenticate v1 and iThenticate 2.0, including running 2.0 directly in the browser and running it through a manuscript tracking system integration, plus a path for upgrading from v1 to 2.0.

Which means a report you are shown may come from one of several interfaces, with different report features available. If someone quotes you a figure from a "Similarity Check report", the version and the exclusion settings behind it are both fair questions, and they are questions that change what the number means.

What an author can actually do with this

  1. Find out whether your target journal participates, and in what. It is a publisher-level fact and often stated in the author guidelines.
  2. Assume full text, not abstracts. Anything you paraphrase closely from a paper you can only see the abstract of is still in the comparison set.
  3. Expect your own prior work to appear, and be ready to say which parts are reused and why.
  4. Ask what the settings were before treating any percentage as a property of your manuscript. It is a property of a comparison, and the comparison has parameters.
  5. Do not treat the number as the finding. The guidance is explicit that a match is not by itself a finding; the useful object is the list of matches and how each one is attributed.

If the report has a second number on it

Some manuscripts come back with an AI writing indicator alongside the similarity figure, where the publisher's licence includes it. That is a separate measurement with a separate meaning, and no amount of citation work moves it.

HumanPen addresses that second number only. Give it the manuscript and the report, and whatever was highlighted is the whole of the job.

For the similarity side, the honest advice is the boring kind. Attribute, quote properly, ask what the exclusion settings were, and be ready to explain reuse. No rewriting product should be sold to you as a solution to a matched reference list.

Frequently asked questions

Can I run Similarity Check on my own manuscript? Not through this route. Crossref describes it as a service for members, providing editors with a tool to check papers; access is tied to member organisations rather than to individual authors.

Does it only compare against open-access papers? No. Crossref describes members as having read-only access to the full text of articles in the Similarity Check database for comparison purposes, and states the database holds over 78 million full-text scholarly content items.

Will my own earlier paper show up as a match? It can. The comparison corpus is built from members' registered content, so previously published work — including yours — is part of what a new manuscript is compared against.

Two editors gave me different percentages for the same manuscript. Which is right? Both, potentially. The search targets are selected per check and the report has exclusion settings, so the percentage describes a particular comparison rather than a fixed property of your file.

KEEP READING