GPTZero Mixed and AI-Polished: What the Labels Mean
GPTZero shows you three percentages, then hides five named classifications one click underneath them. AI Polished is one of the two that live under Mixed. GPTZero defines it in a panel on the result screen, and explains it in a June 2025 post that sits on its news site rather than in its help centre. Here is the whole set, in GPTZero's wording, and what the number beside each one is actually counting.
HumanPen Team
· 14 min read
The short answer
GPTZero prints three percentages, headed "Chance this entire text is...": AI, Mixed and Human. That number is the model's confidence that the class is the right one, not the share of your essay in it. GPTZero says this on the result screen itself, in one line under the chips. Click any chip and it opens into named sub-readings, five in total across the three, and "AI polished" is one of the two under Mixed. GPTZero defines it as "Human-written text that was refined or edited by an AI tool", adding that with this label "the core ideas and structure are the student's own work, but AI was used as an editing assistant". So "100% mixed" is the model saying it is highly confident the document belongs in the mixed class. Which of the two readings under Mixed it leaned on is a separate number, one click away.
What the number beside the label is counting
Start with the screen you are looking at, because GPTZero answers this on it. Expand the Mixed chip and this line appears directly under the readings:
This result isn't the percentage of mixed text, it's the model's confidence in the classification of the entire text as mixed.
The same sentence is waiting under the other two chips, with the class name swapped: "isn't the percentage of AI text" under AI, "isn't the percentage of human text" under Human.
Why the distinction matters is spelled out by Alex Adam, who leads GPTZero's machine learning team, in a post dated 12 January 2024:
categorizing predictions this way enables the confidence of the detector to be disentangled from the percentage of sentences that are AI-generated. Without this distinction, an AI probability of 50% could mean that the detector is 50% confident in its prediction (essentially a coin flip), or that 50% of the text is AI-generated.
Two readings of the same "50%", and they mean opposite things for you. The three-class output exists to keep them apart.
The documentation agrees from three other directions. An entry dated 9 January 2024 in GPTZero's release notes says predictions were "tuned to better represent the probability that the model is correct", giving the worked example that a 90% AI score means "there is a 90% chance that the text is entirely written by an AI". The API help page describes a class probability as "the chance that the detector is correct in its classification", adding that 90% means correct "90% of the time on similar documents". Note similar documents, plural, not yours. And the technology page lists it among GPTZero's research contributions in the driest phrasing available: the detector is framed as "a trinary classification problem, separating prediction confidence from the proportion of LLM text."
One GPTZero page words it the other way round. The FAQ block on the home page says GPTZero "will suggest what percentage of your text is likely to be AI-generated". That sits inside a three-step how-to answer rather than describing a named result field, so it may well be loose marketing copy rather than a contradiction. Either way the working rule is the same: "GPTZero said 30%" does not identify which quantity anyone is discussing, so say which surface and which field it came from. Why the same text scores differently on every detector lines that up against Turnitin's, Originality.ai's and Winston AI's units, which is where the cross-vendor version of this problem lives.
Three chips on top, five named classes underneath
This is the part that catches people, and it is the reason a phrase like "written by a human and polished with AI" seems to come out of nowhere. The top row is three chips. Each one opens.
| Chip | Opens into | GPTZero's line under it |
|---|---|---|
| AI | "chance AI Generated", "chance AI Paraphrased" | "isn't the percentage of AI text" |
| Mixed | "chance AI-Human Mix", "chance AI polished" | "isn't the percentage of mixed text" |
| Human | "Human written" | "isn't the percentage of human text" |
Five named readings in total. "AI polished" is not a fourth bucket beside Mixed. It is one of the two things Mixed can mean. That single fact answers most of the confusion, because a mixed result reads very differently once you know the other reading under it is called AI-Human Mix.
Click the small information icon beside any of those readings and a panel opens, headed "Understanding AI Classifications", with a tab for each of the five. GPTZero's definitions, verbatim:
AI Generated — Text written directly by an AI language model with no meaningful human editing. AI Paraphrased — AI-generated text that was rewritten — either by a tool or by hand — to disguise its origin. AI-Human Mix — Human-written text combined with AI-generated text in the same document. AI Polished — Human-written text that was refined or edited by an AI tool. Human — Text written entirely by a human with no AI involvement.
The AI Polished tab carries a second paragraph, and it is the one worth having in front of you:
The text was originally written by a human, then run through an AI tool to improve grammar, rephrase sentences, or enhance clarity. The core ideas and structure are the student's own work, but AI was used as an editing assistant.
Note where each definition starts from. Polished starts with human writing. Paraphrased starts with AI output. They are not two degrees of the same thing, they are opposite origin stories, and they hang off different chips.
Two smaller things. GPTZero writes this label three ways in three places: "AI polished" on the chip, "AI Polished" in the panel, and "Polished by AI" on the sample buttons on its home page. And the free scan needs a minimum: at zero characters the box reads "Enter at least 250 characters to scan", switching to "Great! You've entered enough text to scan" once you are over.
Where "AI Polished" came from, and where GPTZero explains it
The class is younger than the other four. It arrives in a release note dated 3 June 2025, filed under "Added", as a single line:
Ability to detect "polished" texts that are lightly edited by AI
Nine months later, one more line, 11 March 2026: "Improved performance on the AI polished class." Those two entries are the whole of its appearance in the release notes.
The longer explanation came two days after that first entry, in GPTZero's post introducing the AI Polished class, by Emily Napier, dated 5 June 2025. It says in so many words where the label sits:
To improve transparency, we've decided to distinguish between documents that combine "chunks" of AI and human text and documents rewritten or polished by AI. We now identify these documents under the "Mixed" class with an accompanying tag of "AI Polished".
That is the table above, in GPTZero's own words and with a date on it: two kinds of mixed document, and one of them gets the AI Polished tag.
If you go looking for the post, search for two wordings. Its headline and body use "AI Polished", the panel's wording. Two of its figure captions use "Lightly edited by AI" instead, which is not one of the three above but is the wording of the release note, and the page address is built from it.
The post is not part of GPTZero's help centre, and none of the help articles mentions the label. We checked the support centre exhaustively on 15 September 2026. Its home page lists five top-level collections, each printing its own article count: General 36, API 17, Origin Extension 8, Account Management 11, Microsoft Word Add-In 5. That is 77. Walking into each one gives eight sub-collections plus three articles attached directly to a collection page, and enumerating those yields exactly 77 distinct articles. We fetched all 77 and searched each page's `innerText`, the text the browser renders (non-breaking spaces turned into plain spaces, the site's navigation and sidebar included), 88,894 characters in total that day. The sidebar changes between page loads, so a rerun will not land on exactly that figure:
- `polish`, `polished`, `AI polished`, `AI-polished`, `AI Polished`, `Polished` — 0 each
- `Polish` — 1, and it is the language, in the supported-languages list
- `confidence in the classification`, the phrase the result screen uses — 0
The counter was alive: over the same corpus that day, `burstiness` returned 28, `paraphras` 18, `sentence` 19, `mixed` 11. Paging through the main news index doesn't get you there either. On 15 September 2026 that index was 30 pages deep, ending with 147 distinct post addresses, and the June 2025 post was not one of them. It was listed in the news sitemap, which held 320 posts at the time. GPTZero publishes often, so both totals will have moved by the time you check.
None of this makes the label unreliable, and the count is not an argument that it is. It just explains why reading through GPTZero's help articles leaves you with nothing on AI Polished. If someone asks what your label means, the post is what to send them. Screenshot the definition panel as well as the score, though. It is a short definition, one click from the reading your document got, and it is the difference between a reader who thinks "polished" is a euphemism and one who has read GPTZero saying the core ideas are the writer's own.
Which level the label is actually about
The chips describe the document. The heading above them literally reads "Chance this entire text is...". The sentence highlights are a second output with a different job, and the FAQ is explicit about which way the reasoning runs:
The sentence-level classification should not be solely used to indicate that an essay contains AI (such as ChatGPT plagiarism). Rather, when a document gets a MIXED or AI_ONLY classification, the highlighted sentence will indicate where in the document we believe this occurred.
Document first, highlights second. They explain the verdict rather than add up to it. On screen this appears as a list headed "Your most AI sentences", with each entry tagged for impact rather than given a percentage of its own.
GPTZero also publishes a ranking of how far to trust each level, and it runs against the way most people read a report:
The accuracy of our model increases as more text is submitted to the model. As such, the accuracy of the model on the document-level classification will be greater than the accuracy on the paragraph-level, which is greater than the accuracy on the sentence level.
The individual highlighted sentence is the thing you end up staring at. It is also the layer GPTZero rates least accurate.
Two more behaviours catch people out. The highlighted sentences, says the Advanced Scan help article, are the ones "disproportionately affecting your overall AI or human score", which is impact rather than proportion, and GPTZero adds that unhighlighted sentences "may still have an effect on your score, but they aren't more likely than other parts of your document". Separately, a middling result with no highlights at all has its own help article: "there isn't one specific section of the document that is especially likely to have AI content. Rather, the entire document has signs of being AI generated, but with lower certainty."
Why the number moved when you barely touched anything
Two release notes speak to this, and both are narrower than people assume.
6 February 2024:
Processing of long documents has been made more consistent. There should be less variation in predictions when cutting out a few sentences from long AI-generated documents
2 May 2025:
Increased output stability when rescanning multiple times
Look at the scope on each. The first is about cutting sentences out of long AI-generated documents. The second is about running the same text through twice. Neither is a statement about editing one sentence of your own draft and scanning again, and we could not find one: across the 77 support articles above, `stabilit`, `rescan` and `re-scan` all return 0. That is the gap in this article, and it has a practical consequence: a GPTZero number only means something next to another GPTZero number if you know when each was taken, so record the date with the score.
The reason dates matter sits on the release-notes page. Counting entry dates at the start of a line in the page's `innerText`, it carries 55 dated entries since May 2023, 10 of them in 2026 up to 15 September, the most recent being 12 August 2026 (Model 4.9b). Two from this year: "More granular mixed text predictions and identification of interleaving mixed texts" in May, and "Robustness to headers; headers are now excluded from our model's input" in August. The result screen prints the model version next to the verdict, which is the cheapest thing on it to record. A scan from July and a scan from last week were not run by the same model.
The confidence band is the other half of the reading, and it is the part that gets dropped when a score is quoted at somebody. The Advanced Scan help article maps the bands onto error rates: "Highly confident: our error rate is less than 2%", "Moderately confident: our error rate would be around 10%", "Low confidence: our error rate would be 14% or higher". The technology page puts the top band at "less than 1%" average error instead, on an internal evaluation set. Both figures are GPTZero's own, on GPTZero's own data, they do not match, and the help article carrying the 2% is stamped as last updated a year ago. GPTZero's ML lead is also candid about the limit on the underlying number: the calibration step means the detector is "not overconfident in its predictions" for the most part, but "achieving perfect calibration is nearly impossible". A class quoted without its confidence band is half a result.
What to do with a label
Whether what you did was allowed is a question for your institution, not for a classifier, and GPTZero says as much more plainly than many of the people quoting it: "these results should not be used to punish students", and "we don't recommend using an AI detection result as the only proof for academic punishment or discipline". Its ML lead goes further, and this one is worth knowing if a single scan is being waved at you:
we do not encourage punitive actions based solely on the results of a single scan. Instead, a history of scan results should ideally be established, serving as stronger evidence of AI usage.
Five things that are in your hands:
- Open the chip before you panic. A bare "Mixed" is the least informative thing on the screen. The two readings underneath it are AI-Human Mix and AI polished, and they describe different situations.
- Record the class, the sub-reading, the confidence band, the model version and the date, in one screenshot. All five are on the same panel. Retyping the number throws away everything that makes it interpretable.
- Screenshot the definition panel too, and keep the link to GPTZero's June 2025 post. The panel is behind the information icon, and a screenshot keeps the definition as your screen showed it. The post is where GPTZero explains, in writing, that AI Polished is a tag under Mixed.
- Find out what your institution actually runs before you argue with a number from somewhere else. A self-check on a free tool is not the report your department will be looking at, and the class names, wording and thresholds are not shared between vendors. Two tools disagreeing on your file is the expected outcome, not a malfunction.
- Keep your drafts and your revision history. GPTZero's own advice to educators is to ask for "drafts, revision histories, or brainstorming notes", and it singles out editor version history as what tends to settle the question. That evidence has to already exist by the time you need it, which makes it a today job rather than an after-the-email job.
The output is a reading of the text sitting in front of it, not a record of how that text came to exist. Whichever chip you landed on, the argument you can actually win is about your process, and the evidence for that lives in your drafts.
Further reading
- GPTZero vs Turnitin: Why Scores Won't Match and What Each Tool Actually Measures — for when the two tools return two different numbers on the same file.
- Turnitin removed the purple AI-paraphrased highlight — Turnitin's own label set changed in August 2026, and an old screenshot is not out of date in the way people assume.
- Why the same text scores differently on every detector
- What AI detectors measure beyond perplexity and burstiness — GPTZero says it stopped using both in autumn 2023, when it moved to a deep-learning architecture.
- Flagged, but you wrote it yourself: preparing a response
- Personal Statement Flagged as AI: The Prompt Chose the Shape First
Everything quoted above was read from GPTZero's own pages and from a free scan run while logged out, on 15 and 17 September 2026, and each quotation names where it sits. Where two of those pages give different numbers for the same thing, both are here.
KEEP READING