Why AI Detectors Flag Well-Written Essays
The question assumes the detector read your essay and marked it down for reading too well. Nothing the vendor publishes about the tool has anywhere to put "well written", which changes what a flag can be telling you.
HumanPen Team
· 26 min read
The short answer
It is not grading the writing. Across the three Turnitin help pages we searched, "quality", "style", "polished", "well written", "vocabulary" and "clarity" were absent. The sensitivity control "300 words" was present on all three pages, so the search path was reading page text rather than returning blanks. The FAQ does not frame the detector as a measure of how good the writing is. What Turnitin does publish is a non-quantified list of text that false positives can include. Careful academic prose can resemble two items on that list when it repeats fixed terms or uses parallel structures, but Turnitin does not say how often those properties appear among errors.
That does not prove polish is harmless, and we are not going to pretend it does. What a company documents is the part it chose to document, and a trained model contains a great deal more than its help pages do. It establishes something narrower and more useful: "well written" is not a quantity this system reports on, so a flag is not a verdict on it, and the advice people build on the assumption that it is, which always ends at writing below your own level, has nothing published behind it.
Counting the vendor's own vocabulary
We searched the rendered text of three pages retrieved on 18 August 2026 because `guides.turnitin.com` returned 403 to a plain command-line fetch: "Using the AI Writing Report", "Turnitin's AI writing detection capabilities FAQs", and "File requirements for an AI Writing Report". Those pages set out the mechanism and file rules; release notes and known-issues pages were outside this search. The table records case-insensitive presence or absence rather than exact counts, which can drift as live pages change.
| Search term | Report guide | Detection FAQs | File requirements |
|---|---|---|---|
| `quality` | Absent | Absent | Absent |
| `style` | Absent | Absent | Absent |
| `well written` / `well-written` | Absent | Absent | Absent |
| `polished` | Absent | Absent | Absent |
| `genre` | Absent | Absent | Absent |
| `discipline` | Absent | Absent | Absent |
| `false positive` (secondary control) | Present | Present | Absent |
| `300 words` (sensitivity control) | Present | Present | Present |
The bottom row is the reason the rest of the table means anything: an empty result and a broken retrieval leave the same trace, so a term known to sit on all three pages has to be present first. On the FAQ, "writing" and "segment" are present, while "vocabulary", "word choice" and "clarity" are absent. "Breakdown" is present on the report guide, but only as the name of an interface panel, not as a statistic.
Then there are the pictures, which a text search cannot see at all. The FAQ article itself contains no images; the site logo belongs to the page chrome, not the article. The report guide carries screenshots with no alt text, so we opened them separately. They include a sample report showing 56% AI beside 23% similarity, a page of highlighted prose, and status messages. None uses any of the words above. One does carry a sentence that appears nowhere in the page's own text:
"AI detection includes the possibility of false positives. Although some text in this submission is likely AI generated, scores below the 20% threshold are not surfaced because they have a higher likelihood of false positives."
That is the product declining to print a number in the band where it trusts itself least, and it is worth holding on to, because everywhere else in this material the error rate is described rather than acted on.
What it does have words for
Something is doing the work, and the FAQ names it under the heading "What parameters or flags does Turnitin's model take into account when detecting AI writing?":
"Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers."
The same answer is careful to add that the model "is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions", and that it "learns statistical patterns from our training data" instead. Both sentences are on the page and reading either without the other gets you somewhere wrong: the vendor is denying two named public metrics, not denying that word probability is the currency. What that means for the terms people argue about is laid out in what AI detectors measure.
Now follow the published units upward rather than downward. Words, then sentences, then this:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1 ... Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."
The chain stops there. The largest object that is ever classified is a segment, a handful of sentences wide, and the document number is an aggregation of those, not a second judgement made after reading the whole thing. Almost everything that makes an essay good is bigger than a segment: an argument that holds for eight pages, evidence chosen well, an objection raised in section two and answered in section five. None of that is an object at the scale where classification happens. It is not that the system weighs your reasoning and dislikes it. There is no slot.
What "human" is calibrated against
When discussing false positives in academic prose, the vendor points to a pre-ChatGPT test corpus:
"To bolster our testing framework and diagnose statistical trends of false positives, before every update or new model release, we perform tests on over 700,000 additional academic papers that were written before the release of ChatGPT to further validate our less than 1% false positive rate."
The FAQ asks itself the same thing again under its own heading, "How does Turnitin ensure that the false positive rate for a document remains less than 1%?", and repeats the figure with the word "over" dropped: "we perform additional tests on 700,000 additional academic papers that were written before the release of ChatGPT". Same corpus, two slightly different sentences on one page.
Read how Turnitin describes that reference population: academic papers selected because they were written before ChatGPT's release. That is a date-based description, not a statement that every paper was independently verified as human-authored. The narrower claim is still useful: Turnitin says it re-runs this pre-ChatGPT academic corpus before model releases to watch the false positive rate. The page does not describe paper-by-paper authorship verification.
Which does not settle it, and we would be overclaiming if we said it did. The vendor publishes the target and not the distribution behind it, and the target itself carries a limit that its own next sentence quietly drops. The full version of that argument, including what happens when you multiply a small rate by a real university's annual volume, is in what a 1% false positive rate means.
The error distribution is not published
The page states a false-positive rate of less than 1%, but it does not publish how errors are distributed across the test corpus. It gives no subgroup rates and does not show whether any kind of text accounts for more errors than another.
What the page does publish is a non-quantified list of text that false positives can include, followed by advice for reading the percentage:
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas. If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
Read what is on that list and what is not. Not sophisticated vocabulary, not long sentences, not a strong argument. It names three possible text properties, but gives no frequency for any of them. The third describes a process rather than anything you would have produced writing on your own.
The first two are worth comparing with ordinary academic practice, and what follows is our reading, not the vendor's. If your field asks you to call one thing by one name every time it appears, you may repeat the same term on purpose because a synonym would introduce ambiguity. If you set comparable items in parallel sentences so a reader can hold them side by side, you may reduce structural variation for the same reason. Neither choice is a lapse. Turnitin says false positives can include these properties. It does not say how often, that errors are concentrated there, or that the properties correlate with errors.
What we are not saying, because it is a much larger claim than any of this supports, is that polished writing gets flagged. That sentence circulates widely, and what it usually arrives attached to is an instruction to sound less capable than you are. The edits the published list actually supports, and what doing them by hand costs, are worked through in how to humanize AI text without a tool.
Notice too who that second sentence is written for. It asks whoever is looking at the percentage to weigh what kind of text produced it. The instruction is pointed at the person holding the report, not at the person who wrote the paper.
Whether your kind of writing is worse off is not published
The obvious next question is whether false positives are more common in your discipline, your section, or your genre. The FAQ has a heading that comes close, "Does the Turnitin model take into account that AI writing detection technology might be biased against particular subject-areas or second language writers?", and the answer is about the training sample:
"One of the guiding principles of our company and of our AI team has been to minimize the risk of harm to students, especially those disadvantaged or disenfranchised by the history and structure of our society. Hence, while creating our sample dataset, we took into account statistically under-represented groups like second-language learners, English users from non-English speaking countries, students at colleges and universities with diverse enrollments and less common subject areas such as anthropology, geology, sociology, and others."
The subject names "anthropology", "geology" and "sociology" appear only in paragraphs describing how the sample was built. "Discipline", "genre" and "subgroup" are absent from all three pages. So subject areas appear as an input to training, not as a published testing breakdown. There is no published false-positive rate for anthropology, and the vendor puts no numbers behind the bias answer, which is why it has to be read as a statement of intent.
That gap is why the figure everyone reaches for on this question comes from outside the company: a 2023 study on 91 TOEFL essays, whose scope, controlled edits and real limits we went through in why non-native English writers get flagged by AI detectors more often. It is also why two detectors can hand the same paragraph two different verdicts, which is a separate mess covered in why detectors disagree.
Some of your best work is not in the measurement at all
There is one more reason "it flagged my good essay" is the wrong frame, and it is about what the percentage is a percentage of. Only some of your document is eligible to be scored:
"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures. This percentage is not necessarily the percentage of the entire submission."
Two changes announced in August 2023 moved the boundary further. One says long-form prose inside tables is now processed; the other says a bug that sometimes highlighted AI writing inside a bibliography was fixed and bibliographies are now excluded from processing. Both of those release notes end the same way, telling you to resubmit an existing paper before either change applies to it.
So the bibliography you spent two evenings aligning and inclusion criteria written as short, non-sentence bullets sit outside the scored set. Tables are conditional: long-form prose inside them can be processed, while numeric cells, fragments, and other non-sentence table content may remain outside qualifying prose. The percentage therefore need not describe the whole submission. That is also the mechanical reason it can disagree with the amount of coloured text in front of you, which we take apart line by line in how to read a Turnitin AI writing report.
What to do with all this
Very little of it is about how you write, and most of it is about how much weight the number can carry.
- Sounding less capable is the one move to rule out. Everything above argues against it, and the trade is lopsided in an ordinary way: an examiner's marks land for certain, while the effect on the score is something nobody has measured and you are usually not allowed to watch.
- Ask which passages, not what number. The highlights tell you where the classifier landed. The percentage tells you how much of the qualifying text it landed on, which is a different quantity from how much of your essay.
- Expect a short document to be measured coarsely. The vendor's words: "In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap." The sentence after it is the one that matters: "This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated." Below 300 words of qualifying text, the file requirements say there is no report to read at all.
- If your prose is conventional because the genre demands it, check who the vendor's advice is written for. That sentence is addressed to whoever reads the percentage, not to you. Worth knowing before you conclude that the paper is the thing that has to change.
- Distrust any promise about the resulting number, ours included. Turnitin's own position is that individual predictions "may not always be explainable in simple feature-by-feature terms". Nobody outside the company can map an edit onto a movement in the score.
Where a rewrite fits, and where it does not
None of this is an argument for rewriting. Where the work is yours and you are prepared to say so, the paper is not the object that should be changing.
A rewrite earns its place at one specific point: when a report exists, particular passages are marked on it, and those passages are going to be revised anyway. That is the shape HumanPen is built for. Hand it the file plus the AI report from Turnitin or iThenticate and the marked passages become the scope, with everything outside them keeping the words you wrote. Our own page states the granularity rule: "A paragraph is the smallest unit the engine rewrites: if your selection covers only part of one, it is expanded to the full paragraph and shown that way for you to confirm." Billing counts the words actually rewritten, so a scoped job costs what the scope costs rather than what the file weighs.
If a new report still flags passages, eligible results can continue lowering AI for free. What nobody can hand you is the figure on the next report. After everything above, a tool that quotes you one is quoting a number it does not have.
Frequently asked questions
Is well-written prose more likely to be flagged? That is a bigger claim than the published material supports, and we are not making it. What is documented is a list of three kinds of text that Turnitin says false positives can include, none of which is a quality judgement, plus a measured effect on writers with a smaller lexical range in English from one 2023 sample. Treat the broad version as a claim in circulation rather than a finding.
Does using advanced vocabulary raise my AI score? Nothing the vendor publishes speaks to it: "vocabulary" and "word choice" were both absent from the three help pages we searched. One 2023 controlled experiment moved word choice directly and pushed against the folk theory rather than with it: the researchers enriched the wording of already-flagged essays, and the flags fell. That establishes that word choice was not inert in that experiment. It establishes nothing about your document.
My writing repeats itself because the terminology has to be consistent. That is the conventional case, and it answers to one item on the vendor's own false-positive list. Its advice in that situation is aimed at the person reading the report: take the nature of the text into consideration when looking at the percentage.
Should I add sentence-length variation on purpose? Not as a formatting exercise. Where variation was already carrying meaning, restoring it is ordinary revision. Manufacturing it in a section whose conventions require uniformity replaces writing your discipline asked for with writing it did not, in exchange for an effect nobody can measure from outside.
Can I check the score myself before I submit? Usually not. Turnitin's position is that the AI writing indicator and report are for instructors and administrators, and it reaches a student only if an instructor downloads the report and passes it on. Any advice that assumes you can watch the number move is assuming access most students do not have.
KEEP READING