Turnitin flagged my whole methodology section

The question worth asking is not why the model thinks you cheated in chapter three. It is why the highlighting arrived as one continuous block, and what the percentage attached to it is a percentage of.

HumanPen Team

· 23 min read

The short answer

Because the score does not reach your sentences one sentence at a time. Turnitin's published description is that a submission is cut into overlapping segments, each segment gets a probability, and every qualifying sentence inherits the score of the segments it sits in. Feed that a stretch of writing that holds the same register, the same sentence frames and the same vocabulary for two thousand words, and there is not much for neighbouring segments to disagree about. Highlighting then tends to arrive as a band rather than as a scatter. Methods and materials is the section written that way on purpose, and Turnitin's own account of where its false positives cluster names two properties the genre is required by convention to have. That does not mean the number is wrong. It does mean a wall of red across one section is not forty separate verdicts about forty sentences, and reading it as if it were will send you off to fix the wrong thing.

Why it arrives as a block

What people describe is specific and it is always the same shape. The literature review is clean. The discussion is clean. Then the methods chapter opens and the highlighting starts at the first sentence under the heading and does not stop until the chapter does. That looks less like a model finding suspicious passages and more like a model finding a suspicious chapter, which is why it is so unsettling to look at.

Turnitin publishes how a score gets to a sentence, and it is worth reading the two sentences in the middle of that description rather than the summary of it:

"Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score."

So the unit being classified is a window of text, not a sentence. A sentence is scored by the company it keeps. What follows from that is our reading rather than something Turnitin publishes: where consecutive sentences share windows and the windows look alike, the scores those windows produce have little reason to diverge, and highlighting derived from them comes out contiguous. A band is the expected rendering of a uniform stretch. It is not a count of independent findings.

That has a practical edge to it. The first highlighted sentence in your methods chapter is not necessarily where anything began, because a segment straddling the heading contains the last paragraph of whatever came before. People spend real time staring at the first flagged sentence trying to work out what is wrong with it specifically. Under the published description there may be nothing wrong with it specifically.

It also means the arithmetic people reach for does not exist. You cannot divide a band by its sentences and get a per-sentence cost, and you cannot subtract a sentence and predict what happens. We took that apart against the three most-circulated versions of it in if I rewrite one flagged sentence, does the score drop.

The honest limit on all of this: Turnitin has published the shape of the pipeline and not the pooling function, the aggregation, or the segment length. Everything above is what the published description implies about distribution. None of it lets anyone predict a number.

How much of your methods chapter is even being scored

The second thing to establish is the denominator, because two methods chapters containing the same information can present the model with completely different amounts of text. Turnitin only analyses what it calls qualifying text:

"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures." The sentence immediately after it is the one that matters here: "This percentage is not necessarily the percentage of the entire submission."

Now look at how your discipline actually writes this section. A wet-lab protocol arrives as numbered steps, a reagents table, an instrument list and a settings block. A psychology or education methods chapter arrives as continuous paragraphs: participants, materials, procedure, analysis, each one a run of full sentences. Those are two different submissions as far as the analysis is concerned, and the second one is the one that comes back solid.

How the section is laid outWhat enters qualifying textWhat the percentage is then about
Numbered steps, parameter tables, equipment lists, equationsLittle of it, since most is not standard grammatical sentencesMostly the prose elsewhere in the file, not this section
Continuous paragraphs describing participants, materials and procedureEffectively all of itA sample made almost entirely of conventional procedural prose
Long-form prose sitting inside a table cellIt is processed. A release note dated 9 August 2023 announced the model could handle long-form prose text in tables, and added that existing submissions have to be resubmitted before they are reprocessedThe same as ordinary paragraphs, which surprises people who moved text into a table for other reasons

There is a counterintuitive consequence in the middle row. Being excluded from qualifying text is not a discount. It removes the numbered steps and the tables of values, which were never going to look like generated prose, and leaves behind the part of your writing that is most conventional. The narrower the qualifying sample, the more completely the number is a statement about your procedural paragraphs specifically.

One exclusion runs the other way and is worth knowing because it removes a common false alarm: the same 2023 release note records that highlighting inside a bibliography was a bug, that bibliographies are now excluded when the AI writing report is processed, and that submissions with highlighted reference sections need resubmitting to reprocess. If you are looking at an old report with the reference list lit up, that is what you are looking at.

The gap between the number and the amount of highlighting on the page follows from all of this, and Turnitin says so directly rather than leaving it to be inferred. How to read a Turnitin AI writing report works through that gap; can Turnitin check a whole thesis covers what happens at the length limits, which is where chapter-level submissions live.

The genre answers to the vendor's own list

Turnitin publishes an account of what its false positives tend to have in common. Verbatim:

"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."

The sentence after it is not addressed to you at all:

"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."

Two of the three items on that list describe things a methods section is supposed to do. Structural variation is what a reporting checklist spends its time removing: the same frame for each measurement, the same order of subsections, the same construction for every instrument so that two procedures described identically can be compared. Literal repetition is how you signal that two conditions really were handled the same way. This is the only section of a thesis where writing the same sentence shape eleven times is correct, and it is also the section where doing so is expensive to undo. What you can and cannot safely reword there is a separate problem with its own rules, in what you can rephrase in methods and results, and the wider question of whether to adjust your writing at all is in should you change how you write.

If you checked the section on its own, you measured something else

A lot of the panic in this situation comes from a self-check rather than from an institutional report: the chapter gets pasted into something, or submitted on its own to a draft-check assignment, and comes back at a number nobody wants to see.

Turnitin describes short submissions behaving differently in kind, not just in degree:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap." And the sentence that follows: "This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

A methods chapter submitted by itself is a candidate for that zone, and more so than its word count suggests, because the numbered steps and tables come out of the denominator first. There is also a floor underneath: an AI writing report needs at least 300 words of prose in a long-form format, so a short methods section on its own may produce no report at all rather than a low one. And below the display threshold there is nothing to read: results above 0% and under 20% are shown as an asterisk with no number and no highlights, which means a real improvement and no improvement look identical from where you are standing.

What a red section does not settle

It is tempting to take the mechanism above and treat it as a finding in your favour. It is not one. Everything on this page explains what the number is a measurement of. None of it is evidence about who wrote your chapter, and it cuts both ways: a band across your methods section is not proof of anything, and neither is an explanation of why bands happen.

"Our AI writing detection model may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student." The next sentence is where the weight sits: "It takes further scrutiny and human judgment in conjunction with an organization's application of its specific academic policies to determine whether academic misconduct has occurred."
  • A band is a distribution, not a tally. It tells you the model's scores were uniform across that stretch. It does not tell you how confident any of them were.
  • The percentage is about qualifying prose, so it is not comparable between two chapters laid out differently, and adding chapter percentages together produces a number that does not refer to anything.
  • The report is generated by a system its own vendor says can misidentify text, which is a reason for a person to look rather than a reason to dismiss the output.
  • None of this is checkable by you in most setups. The indicator and report are not visible to students, though an instructor can download the report as a PDF and share it.

Do not carry this article into a meeting as an argument. Explaining a classifier to someone who has already read the same help pages tends to look like preparation rather than like an answer, and it moves the conversation onto ground where you are arguing about a model instead of about your work. What actually carries weight is provenance: drafts, the protocol, the analysis file, the dates. Flagged, but you wrote it yourself covers what evidence is worth assembling, the conversation after an AI flag covers what the meeting is procedurally, and version history as evidence covers where the record lives if you have not thought about it before.

If the section is going to be revised anyway

Sometimes the decision has already been made, by a supervisor or by you, and the question becomes how to do it without breaking anything. Methods is the worst section in a thesis to reword casually. Order, timing, dose, equipment, settings, decision rules, analysis population and exclusions are the content, and a fluent sentence that quietly drops one of them is a real defect where a stiff sentence was not.

The cost people underestimate is not the rewriting. It is the re-verification. Every paragraph that gets touched is a paragraph you now have to read against the protocol, the analysis code and the tables. A plan that touches a whole chapter has created a proofreading job the size of a chapter, which is why the scope of the change matters more than the method used to make it.

That is the part HumanPen is built around. Import the Turnitin or iThenticate AI Writing Report for that submission and the passages it marks become the scope: only that text is rewritten, everything else is frozen, and the site's own description of the paragraph rule is that "a paragraph is the smallest unit the engine rewrites: if your selection covers only part of one, it is expanded to the full paragraph and shown that way for you to confirm". Formulas, figures and table structures are not part of the text rewrite, which matters more in this section than anywhere else in the document. Billing counts the words actually rewritten rather than the file you uploaded. Eligible results can continue lowering AI for free.

What none of that is, is a prediction. Turnitin's model changes, the display rules have changed before, and we have no way to tell you what your next report will say. The claims above are about which words get touched and which get billed, which are the two things we can actually control.

Frequently asked questions

Why is my whole methodology section highlighted when I wrote every word of it? Because the classification unit is an overlapping segment, not a sentence, and sentences inherit and pool the scores of the segments containing them. A long stretch written in one register gives adjacent segments little to differentiate, so the output tends to be contiguous. That is a description of how a band forms, not a statement that your particular band is a false positive.

Does a solid block of highlighting mean every sentence in it was judged to be AI? No, and the published description is the reason. A sentence carries a score it inherited from the segments it sits in, possibly more than one, pooled. Reading a band as a list of individual verdicts is the mistake that leads to sentence-by-sentence editing plans.

My methods section is mostly numbered steps and a parameter table. Why is the percentage still high? Those may not be in the analysis at all. Turnitin analyses qualifying text, meaning prose written in standard grammatical sentences, and it states that the percentage is not necessarily the percentage of the entire submission. Excluding the steps and tables leaves the conventional paragraphs behind as the sample the number describes.

Should I add variety to my methods section so it looks less repetitive? That is a question about your discipline's reporting standards before it is a question about a detector, and the repetition in a methods section is usually doing a job. The vendor's own advice about text with these properties is directed at the person reading the report, not at the person who wrote it.

I ran the chapter through a checker on its own and got a very high number. Is that the number my supervisor sees? Not necessarily, because it is a different submission. Turnitin says predictions on documents of a few hundred words come out mostly all or nothing, since there is a single segment with no overlap, and that mixed content in that situation can be flagged as entirely AI-generated. A chapter is also a different denominator from a whole file.

Can I quote the vendor's false-positive list as a defence? It is a published statement about the tool's error modes and it is aimed at whoever reads the percentage, so it is fair to know it exists. It is not evidence about your document, and taking a mechanism into a misconduct conversation as an argument tends to work against you. Bring the drafts and the protocol.

KEEP READING