How to humanize AI text without a tool: the edits that actually change the statistics

Turnitin scores overlapping multi-sentence segments, not single sentences. Here is which manual edits target the properties its own documentation names, and what doing the work by hand actually costs.

HumanPen Team

· 14 min read

The short answer

Two different jobs get called "humanizing AI text". One is trying to push generated text past a check. The other is the situation most people are in: the writing is yours, or it started as a draft you have since rewritten, and it came back flagged anyway. This page is about the second one. Almost everything below is a reason why the first one is a moving target.

Edit whole passages, not individual words. Turnitin cuts a submission into overlapping multi-sentence segments, scores each segment, and lets every qualifying sentence inherit the score of the segments it sits in, so a synonym swap inside one sentence is competing against the paragraph around it. The three properties Turnitin names in its own false positives are text with little structural variation, text that literally repeats itself, and text that has been paraphrased without developing new ideas. That list is the closest thing to a published specification for hand editing.

Nobody, including us, can tell you a number your score will land on afterwards. Most of the advice on this topic is written as if you can measure your own result. You usually cannot, and there is a section below on that.

What the score is a statistic of

The AI writing percentage is not a count of suspicious sentences. Turnitin's FAQ describes the pipeline like this:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

Three practical things fall out of that.

A sentence does not have its own opinion. Its score comes from the blocks of text it belongs to, and it belongs to more than one. So the sentence you are staring at may be carrying a score that was set by its neighbours.

Only prose counts. Turnitin calls this qualifying text and says it analyses "only blocks of text that are written in standard grammatical sentences", not lists, bullet points or other non-sentence structures. It also says the resulting figure "is not necessarily the percentage of the entire submission", and that a document mixing several writing types "would result in a disparity between the percentage and the highlights". If your highlighted text does not look like it adds up to the number you were given, that is the documented reason, not a glitch. There is more on that in how to read a Turnitin AI writing report.

And the unit you can affect is the passage. Not the word.

The one list of features Turnitin has published

Everything else on this page is downstream of a single sentence in that FAQ:

"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."

The next sentence matters as much: "If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated." That is the vendor telling instructors to discount high scores on writing of this kind.

Be clear about what the list is: Turnitin naming the shapes of human writing its own detector gets wrong. That is not a description of how the model works, and it is not a checklist for disguising generated text. It is a description of the position you are in if you wrote the thing and it got flagged anyway. Three named properties, three edits. That is the whole method.

The popular explanation is narrower than it looks

Before the edits, clear this out of the way, because half the advice online is built on it.

The standard story is that detectors measure burstiness and perplexity, and that academic prose scores badly because it is evenly paced. The same FAQ addresses it: the model "is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions". It says instead that the model learns statistical patterns from training data, and that individual predictions may not always be explainable feature by feature.

Careful with how far you take that. The same page also says its classifiers are trained to detect differences in word probability, and perplexity is a measure of word probability. So the accurate version is narrow: Turnitin denies computing those two named scores as explicit rules. It does not deny that word choice statistics matter.

The practical effect is smaller than it sounds. Varying your sentence rhythm is still worth doing, but for the reason in the previous section, not because you are moving a burstiness dial. We went through where that story came from in what AI detectors measure beyond perplexity and burstiness.

Edit 1: rebuild the paragraph instead of repairing the sentence

Open the flagged paragraph, then close the draft and write the paragraph again from your notes, your outline, or the source it came from. Do not edit the existing sentences. Write the passage, then diff it against the old one to check that no fact, number or citation went missing.

Why this one first: segment scoring means the surrounding sentences are part of what produced the score. A paragraph rewritten from source material comes out with a different sentence inventory, a different order of claims, and usually a different length. A paragraph edited word by word comes out the same shape.

Do it in runs of two or three consecutive paragraphs rather than scattering single-sentence fixes across the document. Segments overlap paragraph boundaries, so contiguous work is doing more than the same number of edits sprinkled around.

This is the slow edit. It is also the only one that reliably produces text you can defend as yours, sentence by sentence, if someone asks you to.

Edit 2: put back the structural variation

First item on Turnitin's false-positive list. Generated drafts and, honestly, a lot of tired human drafts converge on the same shape: paragraphs of near-identical length, each opening with a topic sentence, each closing with a "this shows that" wrap-up, each containing three supporting points.

What to change:

  • Let paragraph length follow the job it is doing. A caveat can be one sentence on its own. A worked example can run twelve, and should, if that is what it takes to work the example.
  • Cut the closing summary sentence from paragraphs that do not need one. In a lot of drafts, every third paragraph ends by restating its own opening.
  • Break the topic-sentence reflex: start one paragraph with the exception, one with the number, one already mid-argument.
  • Some sentences should just stop.

One trap. Converting prose into bullet points is not this edit. Bullets are not qualifying text, so pulling prose out of sentences pulls it out of the analysis, which can move the number without anything about your writing having changed. Your marker will notice, and you will have made the paper worse to fix a statistic. Moving text into a table does not even do that much: a release note says Turnitin can now process long-form prose in tables, and that existing submissions have to be resubmitted for that to apply to them.

Edit 3: delete everything your draft says twice

Second item on the list, and the cheapest win on this page.

Search your own document for repeated strings of six or more words. Then check the four places where academic drafts habitually repeat themselves: the abstract against the introduction, the introduction against the first section, every "as mentioned above", and the conclusion against everything. Also check for the paragraph that restates its own first sentence at the end, which is edit 2's problem wearing a different hat.

Delete the second occurrence, not the first. If a point genuinely needs restating in the conclusion, restate it in one clause, not a paragraph.

Done properly this can take a surprising amount out of a section. It also shortens the paper, which is usually a separate problem you already have.

Edit 4: develop the idea instead of rephrasing it

Third item: "text that has been paraphrased without developing new ideas". This is the only edit on the list that changes what the paragraph claims rather than how it reads, and it is the one that survives contact with an actual marker.

For each flagged passage, add the thing that only you can add:

  • The number from your own data, and the one case where it did not hold.
  • Why you chose this method after rejecting the obvious one.
  • A limitation you noticed while doing the work, rather than one you found listed in somebody else's discussion section.
  • Where two of your sources disagree, and which one you are following.
  • The result you expected and did not get.

If a paragraph cannot absorb any of that, ask whether it is carrying an argument at all. Paragraphs that are true but empty are the ones that read as generated, and by Turnitin's own account they are also the ones its detector gets wrong.

Cost: this one is open-ended. It can mean going back to your sources for an afternoon. It is also the only edit here that improves the mark independently of any detection score.

Edit 5: check whether your document is even long enough for this to work

Turnitin says short documents behave differently:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

And the next sentence: "This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

For a 500-word personal statement or a discussion post, there is effectively one segment. Editing two sentences inside it is not partial progress towards anything. Either rewrite the piece whole from your notes, or accept that the score is behaving as documented and put your effort into being able to show how you wrote it. We covered that side in flagged, but you wrote it yourself.

What each edit targets, and roughly what it costs

The effort column is an estimate, not a measurement. It assumes English academic prose, your notes to hand, and no interruptions, which is a set of assumptions that has never once been true.

EditWhat it targetsRough effort per 1,000 words
Rebuild paragraphs from your notesSegment-level scoring: the passage is the unit, not the sentence45 to 90 min
Restore structural variation"content without a lot of structural variation"15 to 30 min
Cut literal repetition"text that literally repeats itself"10 to 20 min
Develop the ideas, add what only you know"paraphrased without developing new ideas"Open-ended, often hours
Convert prose to bulletsRemoves prose from the analysis rather than changing itDon't

Read the first four rows as one job, not a menu. They overlap heavily: rebuilding a paragraph from notes usually fixes its length, its repetition and its emptiness in the same pass.

You probably cannot measure any of this

This is the part most guides skip, and it changes what you should do.

The AI writing indicator and report are not visible to students. Only instructors and administrators see them, though an instructor can download the report as a PDF and share it. So unless someone shows you, you are editing without a gauge.

Even with the report in front of you, the resolution is poor at the bottom. Turnitin does not display a percentage for scores between 1 and 19, only an asterisk, and says it does this to avoid potential false positives. An edit that took a document from 18% to 6% therefore shows up on the report exactly the same way as an edit that did nothing.

The model is also built to under-report. Turnitin says that to hold its false positive rate down, "there is a chance that we might miss some AI written text in a document", and gives its own example: a document reported at 50% "could contain as much as 65% AI writing". A number that moves down is not proof of much in either direction.

One thing worth holding on to while you do this work. Turnitin's guidance to educators is that the score is being used as intended when it is treated as "a single data point rather than a definitive response", and three of its help pages say the score should not be the sole basis for adverse action against a student. Fix the paper. The score sits downstream of the paper, and you cannot see it anyway.

What doing this by hand actually costs

Time is the obvious one and it does not scale linearly. A 1,500-word essay is an evening. An 8,000-word chapter with tables, a reference list and cross-references is a different animal, because the editing is no longer the hard part.

The less obvious costs:

  • Every paragraph you rewrite is a paragraph you have to verify again. Numbers, tense in the methods section, whether the citation still points at the claim it was attached to. The verification is often longer than the rewrite.
  • You can flatten your own argument. Three passes of "make this less uniform" on your own prose, at 1am, is how a careful hedge turns into an overstatement.
  • On long files the damage is structural. Cross-references, numbered captions, footnotes and TOC fields break when text moves, and they break silently. We wrote about that failure mode in Word fields, tables of contents and cross-references.
  • You may redo work that was already fine, because you cannot see which passages were flagged unless your instructor shares the report.

Scope is the expensive part, and it is the part HumanPen was built around, so treat this paragraph as an interested party talking. The report-import path takes the AI Writing Report you were given, and in the words of that page, "Only passages matched from the report are rewritten; the rest of the document is preserved verbatim". Eligible results can continue lowering AI for free. Do the edits by hand or don't. The principle is the same either way, and it is the boring one: the narrower the scope, the less there is to re-verify afterwards.

When to do it by hand anyway

Manual editing is the right call more often than a tool vendor would like to admit:

  • The document is short. On a few hundred words you are rewriting the whole thing regardless.
  • You still have your notes and sources open. Edit 4 is nearly free at that point and impossible six months later.
  • The flagged section is one you can genuinely rewrite from the underlying work, like a methods section or a results paragraph.
  • The problem the report found is a real problem. Repetition, empty paragraphs and uniform structure are writing faults first. Fixing them improves the paper whether or not any number moves.

Frequently asked questions

Does varying sentence length lower an AI detection score? Not for the reason usually given. Turnitin says its model is not explicitly programmed to evaluate signals such as burstiness or perplexity. But it separately names "content without a lot of structural variation" as a property common in its own false positives, so variation is worth restoring. Treat it as one edit among several, not the trick.

How much do I have to change before the number moves? Nobody can tell you, and anyone who gives you a percentage is making it up. Turnitin scores overlapping segments and pools the results, so the effect of an edit depends on the passage around it. You also lose the feedback in the 1% to 19% band, where the report shows an asterisk instead of a number.

My reference list was highlighted. Should I edit my references? Check which report you are holding first. A release note says bibliographies are excluded when the AI writing report is processed (existing submissions only pick that change up on resubmission), and Turnitin says AI writing highlights are not visible in the Similarity Report at all. A highlighted bibliography is almost always a similarity match, and matching reference entries are the ordinary case, since a citation style exists to make everyone format the same source the same way.

I only used a grammar checker. Does that get flagged? Turnitin's FAQ says that in tests on human-written documents, "in most cases", changes made by Grammarly and other grammar-checking tools were not flagged by its detector. It then says this excludes Grammarly's generative features, naming draft generation, paraphrasing and summarizing, and that content produced with those features "will likely be flagged as AI-generated by our detector". The line it draws is between correcting your sentences and generating them.

Is there a score I should be aiming for? No published threshold exists, and institutions set their own practice. Turnitin's guidance to educators is that the score is being used as intended when it is treated as a single data point rather than a definitive response. We went into why the popular 20% figure is not a rule in is 20% AI too high.

Sources: Turnitin's [AI writing detection capabilities FAQs](https://guides.turnitin.com/hc/en-us/articles/28477544839821-Turnitin-s-AI-writing-detection-capabilities-FAQs), [Using the AI Writing Report](https://guides.turnitin.com/hc/en-us/articles/22774058814093-Using-the-AI-Writing-Report), and the [AI writing detection model release notes](https://guides.turnitin.com/hc/en-us/articles/28294949544717-AI-writing-detection-model). Quotes checked against those pages on 14 August 2026. Published on the blog of a company that sells a document humanizer, which is worth knowing when you read the section on what manual editing costs.

KEEP READING