Is ChatGPT an AI detector? What OpenAI published about its own classifier
A growing number of accusations start with a screenshot of a chatbot saying "this was likely written by AI". It is worth knowing what the company behind that chatbot has published about doing exactly this.
HumanPen Team
· 7 min read
The short answer
No. OpenAI built a separate, purpose-trained classifier for this, published its measured accuracy — it correctly identified 26% of AI-written text while wrongly labelling 9% of human-written text as AI — and took it offline on 20 July 2023, stating it was withdrawn "due to its low rate of accuracy". A chat model asked "did AI write this?" is not that classifier and never was; it produces a plausible-sounding answer because producing plausible-sounding answers is what it does.
What OpenAI actually built, and what it measured
In January 2023 OpenAI released a classifier "trained to distinguish between text written by a human and text written by AIs from a variety of providers". The announcement is still online, with a withdrawal note added at the top.
The numbers are in the announcement itself, not in someone's independent test:
"In our evaluations on a 'challenge set' of English texts, our classifier correctly identifies 26% of AI-written text (true positives) as 'likely AI-written,' while incorrectly labeling human-written text as AI-written 9% of the time (false positives)."
Twenty-six per cent caught. Nine per cent of human writing wrongly flagged. And that is the vendor's own evaluation of a system it built on purpose, with access to its own models' outputs as training data.
The opening paragraph contains a sentence that gets quoted less often than it should: "it is impossible to reliably detect all AI-written text."
The limitations it published about itself
The announcement carries a limitations section, and it reads like it was written by people who expected to be misused. In its own words:
- "It should not be used as a primary decision-making tool, but instead as a complement to other methods of determining the source of a piece of text."
- "The classifier is very unreliable on short texts (below 1,000 characters). Even longer texts are sometimes incorrectly labeled."
- "Sometimes human-written text will be incorrectly but confidently labeled as AI-written by our classifier."
- "We recommend using the classifier only for English text. It performs significantly worse in other languages and it is unreliable on code."
- "Text that is very predictable cannot be reliably identified." The example given is a list of the first 1,000 prime numbers, "because the correct answer is always the same".
- "Classifiers based on neural networks are known to be poorly calibrated outside of their training data. For inputs that are very different from text in our training set, the classifier is sometimes extremely confident in a wrong prediction."
That third bullet and the last one are the same phenomenon stated twice: confidence and correctness are not connected. A tool of this kind saying "definitely AI" is not a stronger claim than it saying "possibly AI". It is the same machinery with a different number attached.
Then it was taken down
At the top of the same page:
"As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy."
A company withdrawing its own product and naming accuracy as the reason is about as clean a piece of evidence as this field produces. It costs the company something to say, which is exactly why it is worth more than a third-party opinion in either direction.
The chat model is not the classifier
This is the part that most needs saying, because the withdrawal is sometimes reported as though the capability moved into the chatbot. It did not. The classifier was a separate fine-tuned model, described in the announcement as "a language model fine-tuned on a dataset of pairs of human-written text and AI-written text on the same topic", with a deliberately adjusted confidence threshold "to keep the false positive rate low".
None of that describes a chat interface. When a chat model is asked whether a passage was AI-written, it has no detection component to consult; it generates the kind of text that tends to follow that kind of question. It will produce a confident-sounding verdict, a percentage if you ask for one, and a list of reasons — and it will often produce a different verdict for the same passage on a second attempt.
So the screenshot in circulation is not a weak measurement. It is not a measurement.
What this does and does not say about Turnitin
One vendor's withdrawn classifier says nothing about a different vendor's current product, and it would be dishonest to stretch it that far.
Turnitin's detector is a different system, described by its maker in different terms, with its own published claims and its own documented failure modes. Detectors in general disagree with each other substantially on the same passage, which is a separate problem with its own evidence — set out in how far apart detectors land on identical text. And the population most affected by false positives is documented, in a way that matters if English is not your first language: see what the research shows about non-native English writing.
The narrow claim here is the only one the source supports: asking a chat model whether something was AI-written is not detection, and the company that makes the best-known chat model published a purpose-built classifier, measured it, and withdrew it.
If this is how you were flagged
Some practical shape, without pretending we know your institution's process.
- Establish what produced the claim. "It was flagged" can mean an institutional detector with a report, or a person pasting your work into a chatbot. Those are different situations and only one of them comes with a document.
- If there is a report, ask for it. A report has passages and locations in it; a screenshot of a chat window has a sentence.
- Ask a chat model the same question twice, on the same passage, in two fresh sessions. Instability is a property you can demonstrate rather than assert.
- Do not lead with any of this. Process and provenance persuade; arguments about tooling rarely do, and going first with a technical objection can read as avoiding the question.
Where we sit in this
We rewrite documents, so it is worth being plain: nothing above is an argument that detection does not work, and we do not sell one. HumanPen is for the situation where a real report exists and specific passages carry highlights. Those highlights set the boundary of what it edits; nothing outside them is touched.
If your situation is a chatbot screenshot with no report behind it, you do not have a rewriting problem. You have a conversation to have, and the material above is what we would bring to it.
Frequently asked questions
Can ChatGPT tell whether text was written by AI? OpenAI built a separate classifier for that task rather than using the chat model, published its accuracy, and withdrew it in July 2023 citing its low rate of accuracy. A chat interface has no detection component to consult.
What were the numbers on OpenAI's classifier? On a challenge set of English texts it correctly identified 26% of AI-written text and incorrectly labelled human-written text as AI-written 9% of the time.
Why does the chatbot sound so certain? OpenAI's own limitations note that classifiers of this kind can be "extremely confident in a wrong prediction", and a chat model's confident tone is a property of how it writes rather than evidence about your text. Asking twice often produces two different answers.
Does this mean Turnitin's AI detection does not work either? It says nothing about it. Different vendor, different system, different published claims. The only thing this source supports is that using a chat model as a detector is not detection.
KEEP READING