Randomised Trial Reports and AI Detection: What CONSORT 2025 Decided Before You Wrote It

A randomised trial report is written twice before anyone reads it: once as a protocol, once as a paper, both against numbered checklists from the same group. The published checklist reaches down to a single word in your title. That is worth knowing before you decide which sentences to change because a detector highlighted them.

HumanPen Team

· 15 min read

Why does a randomised trial report come back with a high AI score when the data are mine?

Because most of a trial report is specified in advance and much of it is a second telling of something you already wrote. CONSORT 2025 lists thirty items a report has to cover, tells you to put a particular word in the title, and pushes the participant counts into a flow diagram and the baseline characteristics into a table. What is left in continuous prose is the Methods narrative and the Discussion, and the Methods narrative is a condensed retelling of a registered protocol. Turnitin, in documents that never mention clinical trials at all, says it measures only prose sentences and that its own false positives can involve text with little structural variation or text paraphrased without new ideas. A flag on a Methods section is not evidence of anything about how the section was written, and cutting checklist content to make it look less uniform makes the report worse.

Nothing on the vendor's side says any of this about trials. We read the three pages that define the AI Writing Report on 27 August 2026, 38,174 characters of rendered article text between them, and searched: `CONSORT` 0 hits, `clinical` 0, `trial` 0, `randomis` 0, `randomiz` 0, `protocol` 0, `checklist` 0, `registr` 0. The instrument was working on the same pass: `prose` 12 hits, `qualifying` 12.

The other direction is the one that is easy to get wrong, so we counted it as well. Across the thirteen content pages listed on the SPIRIT-CONSORT site map, plus the CONSORT 2025 checklist and its twelve-page expanded version, 61,768 characters in all: `detection` 0 hits, `detector` 0, `Turnitin` 0, `generative` 0, `large language` 0, `ChatGPT` 0, `plagiarism` 0. Sensitivity on the same pass: `consort` 107, `trial` 167, `randomis` 36. `Artificial intelligence` returns exactly one hit, on the extensions page, where it names a SPIRIT extension for trials whose intervention is an AI system. That is a standard for reporting a trial of an AI tool, which is a different subject from an author using one. So the connection drawn below is ours, assembled from two sets of documents that have never referred to each other.

The guideline names a word for your title

Item 1a of the CONSORT 2025 checklist reads `Identification as a randomised trial`. The expanded checklist, which is where the group spells each item out, puts it plainly:

`Use the word "randomised" in the title`

That is a reporting guideline reaching into a single word of a single line. It is a good rule and it exists for a good reason, which is that indexers and systematic reviewers need to find your trial. It also means that the first line of every report in this genre, worldwide, carries the same token.

Item 1b then specifies the abstract. Its list of required contents is seven bullets, three of which open into ten more, and between them they name fourteen things the abstract has to carry: objectives, trial design and framework, eligibility criteria and settings, interventions and comparators, primary outcomes, how participants were allocated, who was blinded, numbers randomised per group, numbers analysed per group, a result per group with effect size and precision, important harms, a general interpretation, the registry name and number, and the funding sources. Underneath the list is one more instruction, printed in asterisks in the source:

`Do not report information that does not appear in the body of the paper`

Read that as a writing constraint and it is unusual. Your abstract is required to be a compression of your own paper and required not to contain anything new. There is a phrase for that on the vendor's side, in the paragraph where Turnitin describes what its false positives tend to look like, and it is `text that has been paraphrased without developing new ideas`. Nobody is claiming those two documents are talking about the same thing. They are not. But an abstract written to item 1b sits in that description by construction, and apart from its section labels it is prose sentences from the first word to the last.

The Methods section is the second time you wrote it

This is the part that has no equivalent in most other genres, and it is the reason a trial report is a different problem from a paper with a checklist stapled to it.

Before the trial ran, there was a protocol. CONSORT item 3 requires you to say where it is: `Where the trial protocol and statistical analysis plan can be accessed`, with the expanded checklist asking for a URL to its location. Item 2 requires the registry name, the identifying number, the URL and the date of registration. Item 10 requires you to report `Important changes to the trial after it commenced including any outcomes or analyses that were not prespecified, with reason`.

Stack those three up and they describe a specific situation. Your eligibility criteria, your randomisation method, your allocation concealment, your blinding procedures and your planned analyses were all written down in a document that is publicly pointed at from your paper, and the paper reports the same content again in shorter form. That is not a shortcut. It is the point of prospective registration.

The protocol side has its own guideline from the same group, SPIRIT 2025, a 34-item checklist published alongside CONSORT and funded by the same grant. The two documents therefore share more than their subject matter. They were written against paired numbered lists that get revised together.

If you have ever felt that your Methods section reads like a form, this is why. Two of the most heavily specified documents in academic publishing are describing the same procedures, one after the other, to two aligned checklists.

The numbers were moved out of the prose on purpose

CONSORT goes past saying what to report. For several of the heaviest items it says what shape to report it in.

CONSORT 2025 itemRequired formWhere that leaves it for a prose-only measurement
22a, participants randomly assigned, receiving intervention, analysed`In a flow diagram, the number of participants:` followed by a bulleted list of countsA diagram of counts holds no prose sentences
22b, losses and exclusions after randomisation with reasonsAlso `In a flow diagram`Same
25, baseline demographic and clinical characteristics`A table showing baseline demographic and clinical characteristics for each group`A grid of means, SDs and percentages
27, all harms or unintended events`For each group preferably in a table with the number of participants at risk`Mostly tabular, with some prose around it
2, trial registrationRegistry name, number, URL, dateShort lines rather than sentences
5a and 5b, funding and conflictsNamed funders, roles, declared interestsOften a declarations block of short entries

Now put Turnitin's definition of what it reads next to that:

`This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures.`

And the sentence that follows it on the same page:

`This percentage is not necessarily the percentage of the entire submission.`

The reference list is out too, by the vendor's own release note from August 2023, which says a bug highlighting AI writing inside bibliography references was fixed and that `Bibliographies are now excluded when processing the AI writing report.` That note adds that a submission has to be resubmitted before the change applies to it.

So subtract the flow diagram, the baseline table, the harms table, the registration line and the reference list. What is left in the measurement is the introduction, the Methods narrative, the results narrative around the tables, and the Discussion. In other words, the parts CONSORT specifies most tightly, minus almost all of the parts that carry your actual numbers. What Is "Qualifying Text" in Turnitin AI Detection? covers the general rule and Does Turnitin Read Text Inside Tables? covers the table case, including the part of that release note about long-form prose inside table cells.

There is a second consequence worth flagging. A short report whose analysed prose comes to only a few hundred words falls into the behaviour Turnitin describes as `mostly "all or nothing" because we're predicting on a single segment without the opportunity to overlap`, and the sentence after that one says `some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated`. A trial report with a big diagram, three tables and forty references is shorter, in the sense being measured, than its page count suggests. Short Documents and the All-or-Nothing Problem in Turnitin works through that.

Where the checklist rules on your wording as well as your content

Two items go further than listing content, and both are easy to miss because they sit as a single line inside a long table.

Item 22b, on losses and exclusions after randomisation, adds this:

`The wording "protocol deviation" is not sufficiently explicit and exact reasons should be reported`

A reporting guideline naming a phrase and refusing it is rare. The reason it matters here has nothing to do with any detector. Replace a specific reason with a tidy summary phrase during a revision and you have broken item 22b. That failure is invisible to a proofreader and obvious to a reviewer who knows the checklist.

Item 29, on interpretation, is the other one:

`Avoid overinterpretation ('spin')`

The Discussion is normally described as the free part of a paper. In this genre it comes with an instruction about how far your claims may go, alongside requirements to relate the results to existing evidence and to balance benefits against harms. Item 30 then names the four things your limitations paragraph has to address: methodological limitations, imprecision, generalisability, and multiplicity of analyses where relevant.

What the checklist never touches

Here is the distinction that decides what you do next, and it is not the one most revision advice offers.

CONSORT specifies what must appear, and in several places the form it must appear in, and in two places a phrase to avoid. It does not write your sentences. Nothing in the checklist tells you how to phrase why the question mattered.

Fixed, and traceable to a number. The useful thing here is not a general warning about what not to touch. It is that each of these has an item number you can point at. The counts in the flow diagram are 22a and 22b. The registry name, number, URL and date are item 2. Eligibility as it was actually applied is 12a. Who generated the allocation sequence and by what method is 17a. Who was blinded and how blinding was achieved are 20a and 20b. Result per group with effect size and precision, and both absolute and relative effects for binary outcomes, is item 26. That changes the shape of a conversation, because "I cannot cut this" lands differently with a citation attached to it. For the general discipline of separating expression from scientific state, across CONSORT, STROBE and PRISMA together, see Methods and results: what you can rephrase and what must stay exact.

Yours, and usually under-used. Item 6 asks for `Importance of the research question` and `Why a new trial is needed in the context of available evidence`, with sub-items covering how the intervention might work and why you chose that comparator. Item 29 asks how your results relate to existing evidence. Those are arguments, and arguments vary in shape because the thinking in them varies. If your introduction could be swapped with the introduction of any other trial in your field, that is habit rather than the guideline, and it is the one place in the report where variation costs you nothing.

The checklist itself does not pretend to cover every design either. Its extensions page lists twenty-two CONSORT extensions built on the 2010 checklist and one built on the 2025 one, and says in its own words that the list `is by no means exhaustive`. Cluster, crossover, stepped-wedge, pilot, pragmatic, non-inferiority, harms, patient-reported outcomes, each with its own additional items. If your trial is one of those, the specified fraction of your paper is larger, not smaller.

If the Methods or the abstract came back highlighted

  1. Look at where the highlighting sits before you look at the number. A Methods section arriving as one continuous block has a mechanical explanation worth understanding first: Turnitin flagged my whole methodology section.
  2. Count the analysed prose, not the pages. Take out the flow diagram, the tables and the reference list, then see how much continuous prose remains. That number, not the page count, is what the all-or-nothing behaviour applies to.
  3. Mark the checklist content before touching any wording. Open the expanded checklist next to your manuscript and mark every sentence that exists because an item required it. Decide that first, cold, not while rewriting.
  4. Bring the vendor's own position to the conversation. Its FAQ says the percentage `should not be used as the sole basis for action or a definitive grading measure by instructors`, and the paragraph on false positives ends by advising the reader to take the nature of the text into account when reading the percentage. What your instructor is told to do when your AI percentage is high collects that guidance.
  5. If somebody quotes a threshold at you, Is 20% AI too high? explains why no published number works as a pass mark.
  6. Check your disclosure separately. Whether you have to declare an AI-assisted revision is a journal question, not a detector question, and Publisher AI disclosure policies, compared sets the major ones side by side.

If the wording is going to change anyway

Revising a trial report has a specific cost that has little to do with writing. Any paragraph you open has to be read back afterwards against three other documents: the populated checklist, the registered protocol, and the statistical analysis plan that item 3 told the world where to find. Open the whole manuscript and you have signed up to re-verify the whole manuscript. A missed check here is not a typo. It is a published claim about what happened to people who agreed to be in a study.

Which is the reason to care about scope. HumanPen is ours, so treat the next few sentences as a first-party description of how the boundary works rather than a recommendation. The scope is drawn before anything runs: you either add the passages yourself or import a detection report and let the marked passages define them. The site puts the rule this way: a paragraph is the smallest unit the engine rewrites. Mark half of one and the boundary grows out to the paragraph edge, and the grown version is what you are shown to approve. Nothing beyond it is opened. Billing follows the same boundary, counting only the words actually rewritten. What comes back is an editable document, with meaning, structure, terminology, citations, layout and styles among the things the rewrite aims to preserve, and the site's own advice is to review complex documents after download. Where a passage qualifies and is still flagged, it can be re-run at no charge. Rewriting only the paragraphs a Turnitin report flagged walks through the report-driven version of that.

What a score will be afterwards is not something anyone can promise in advance, us included. Whatever you use, put the returned Methods paragraphs beside the protocol and the returned effect sizes beside your analysis output before you resubmit. A confidence interval that lost a digit is a worse problem than a highlighted sentence.

Frequently asked questions

Does following CONSORT raise my AI score? Nothing the vendor publishes says so, and in the read described above its three report pages do not mention trials, protocols or reporting guidelines at all. What it does publish is that its false positives can involve content without a lot of structural variation, and CONSORT removes structural variation from parts of a trial report deliberately.

My Methods section repeats my registered protocol. Is that a problem? Not for reporting. The protocol is what makes the trial prospective, item 3 requires you to say where it lives, and item 10 requires you to report any departure from it. It is a problem only in the narrow sense that a second telling of your own document is dense, uniform prose, which is the shape most exposed to the measurement.

Do the flow diagram and the baseline table count toward the percentage? Only qualifying prose sentences are analysed, and a diagram of counts and a grid of means are not prose sentences. A release note from August 2023 does say long-form prose inside table cells is processed, and that existing submissions must be resubmitted to reprocess, so a table full of full sentences is a different case from a table of numbers.

Can I reword the participant flow narrative to lower the score? You can move sentences, but item 22b requires exact reasons for exclusions and explicitly refuses `protocol deviation` as sufficient wording. If a revision replaces a reason with a summary phrase, the report is now less compliant than the one you submitted, and no score offsets that.

My whole abstract is highlighted. An abstract written to item 1b has fourteen required contents, an instruction not to include anything absent from the body, and no tables, lists or references to break it up. It is the densest qualifying text in the document. The vendor's guidance is that the percentage is not a sole basis for action, and its false-positive paragraph is the thing to put in front of whoever is asking.

Does this apply to CONSORT 2010 papers or to observational studies? The items quoted above are CONSORT 2025, published in 2025 with a 30-item checklist. Observational studies follow a different guideline entirely. The general move holds either way: find out which of your sentences a reporting standard specified before deciding which ones to change.

KEEP READING