How Often Do Chatbots Call Human Writing AI? A 2026 Test of 15 Models

Your teacher pasted your essay into a chatbot, asked whether AI wrote it, and got a yes. A preprint posted in September 2026 measured how often the AI models behind chatbots give that answer about text people actually wrote: 15 models from three generations, each asked the same question about the same 1,000 human-written texts. Which model was asked turned out to matter a great deal. Below are the numbers, and what they do and do not tell you about the answer you got.

HumanPen Team

· 9 min read

How often do chatbots call human writing AI?

It depends heavily on the model. In a 2026 preprint, 15 general-purpose AI models of the kind behind chatbots were each asked, through their APIs and with one fixed question, whether 1,000 human-written texts came from a human or a machine. The previous generation (GPT-4o and four other 2024–25 models) answered "machine" for 36.5% of the human texts on average. The newest generation (eight 2026 models, including GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro Preview) did so for 3.6%, with individual models ranging from 0.1% to 8.6%. So how often a chatbot wrongly answers "machine" about human writing depends heavily on the model: in this test, anywhere from more than a third of the time to almost never. A rate like that describes a large set of texts, not your one essay, and a chatbot's answer is not proof on its own: the test texts were not student essays, and a screenshot may not show which model answered or how it was asked.

The study is "Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis" by Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Ruizhe Li and Shujun Li, of the University of Kent, Shanghai Jiao Tong University and the University of Birmingham. It was posted to arXiv on 28 September 2026. It is a preprint and has not been peer reviewed.

What the researchers did

They started with 1,000 human-written texts from M4GT-Bench, a public research dataset, which the paper describes as "covering a broad range of topics, styles, and formats". We paged through the released file and found encyclopedia entries, how-to guides, forum answers, research abstracts and peer reviews of research papers. The paper does not describe any of them as student coursework. Each of the 15 models also wrote 1,000 texts from matching prompts, most of which asked for 150 to 300 words.

Then every model took a turn as the detector. Each text went to each model with the same instruction, which opens "Please determine whether the following text is generated by large language models or by a human", and the model had to answer "human" or "machine" and explain itself. That produced over 233,000 valid judgements.

The researchers grouped the models by release date:

  • Oldest (2023 to early 2024): GPT-3.5 Turbo Instruct and Mixtral 8×22B Instruct.
  • Previous (2024–25): GPT-4o, DeepSeek-V3, Llama 3.3 70B Instruct, GLM-4 32B and a 72-billion-parameter Qwen model (listed as Qwen2.5-72B Instruct in the paper's model table and as Qwen2-72B elsewhere in it).
  • Newest (2026): GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro Preview, DeepSeek-V4-Pro, Qwen3.6-Max-Preview, GLM-5.1, MiMo-V2.5-Pro and Kimi K2.6.

The models were reached mainly through OpenRouter and the vendors' own APIs, and the paper does not describe testing the ChatGPT, Gemini or Claude apps. Nor does it give dates for the runs: it calls the newest group the most recent models "at the time of our experiments" and records checking the models' default settings on 3 September 2026.

Our earlier article Is ChatGPT an AI detector? What OpenAI published about its own classifier is built around the classifier OpenAI made, measured and withdrew in 2023, and argues from it that asking a chat model whether AI wrote something is not detection. That article does not cite measurements of chat models doing this job. This study adds them: 15 chat models, tested directly on the question, under API conditions that are closer to a teacher pasting an essay into a chat window than a separate classifier is, though still not the same thing.

What happened to the human texts

Models acting as the detectorHuman texts labelled "machine"AI texts labelled "human"
Oldest (2 models)13.0%81.9% to 84.1%
Previous (5 models)36.5%39.6% to 48.1%
Newest (8 models)3.6%6.1% to 13.4%

Each figure is the average across the models in that group. The range in the last column depends on which generation wrote the AI texts.

The three groups went wrong in different directions. The oldest models mostly answered "human", so they wrongly flagged fewer people than the next generation did and let more than four in five AI texts through. The previous generation swung the other way: it caught more AI text but called more than a third of the human texts machine-written. The newest group was low on both counts, although it missed more of the AI text written by 2026 models (13.4%) than of older AI text.

That first row is a trap if you read it alone. A 13% false-alarm rate looks better than 36.5%, but it came from models that rarely said "machine" at all. A detector that never accuses anyone also never catches anyone.

Inside the newest group, the model still mattered

The 3.6% average hides a wide spread. The authors' released results give each newest-generation model's rate on the same 1,000 human texts:

Newest-generation modelHuman texts labelled "machine"
Gemini 3.1 Pro Preview0.1%
Kimi K2.61.2%
Qwen3.6-Max-Preview1.3%
GLM-5.11.4%
GPT-5.52.0%
Claude Opus 4.76.1%
MiMo-V2.5-Pro7.9%
DeepSeek-V4-Pro8.6%

Across all 15 models the rate ran from 0.1% to 55.3%, the highest being Llama 3.3 70B from the previous generation. The paper's headline finding is broader: how well detection works "is primarily driven by detector capability rather than generator provenance". Which model you ask matters more than which model wrote the text.

This table shows how much models differ. It is not a lookup for your case: a chat app may run a different version from the one tested, your teacher's question was probably worded differently, and each model answered once per text.

What the chatbot's reasons are worth

Screenshots usually come with reasons: too structured, too generic, too smooth. In this study every model had to give one. The researchers then had another AI model, GPT-5.6 Sol, compare a sample of those explanations side by side, and it found the same features used to argue both ways. Detailed, coherent writing, the authors write, "may be treated as evidence of human expertise by one detector, but as evidence of formulaic LLM generation by another."

They also caution that the explanations "may be post-hoc justifications rather than accurate and reliable descriptions of how a decision was made". A keyword count the authors ran over the same explanations pointed the same way, and they say human annotators still need to confirm these findings. If the screenshot says your essay has clear structure and smooth transitions, that describes your essay. It does not show where the essay came from.

What this study can't tell you

  • How chatbots treat student essays. None of the human texts were described as student work, and both sides were short: the human texts had a median of 233 words by our count, and the AI texts were mostly asked for 150 to 300. A long essay could come out differently, and the study gives no reason to guess in which direction.
  • Whether asking again changes the answer. Each model judged each text once, so the study says nothing about how often the same chatbot flips on a second try.
  • Whether it holds up. It is a preprint, and the authors themselves call their three-generation split "rather simplistic".
  • Next year's models. Many of the models the researchers first considered had already been deprecated or discontinued, which is partly why the oldest group has only two. The figures describe these 15 models, not whatever comes next.

If a chatbot said your essay was AI

  1. Find out exactly what was run. Ask, politely: which chatbot was it, could you see the whole conversation including the question that was typed, and is this the only check or is there also a Turnitin or other report? The answer changes what you are responding to.
  2. Read your course and institution rules. Look for how suspected AI use is handled and what counts as evidence. Some universities restrict which detection tools staff may use or how much weight detector output carries; 19 Universities That Limit AI Detector Use or Evidence: Official Statements collects the official wording. Whether a chatbot falls under those rules depends on how yours are written.
  3. Bring your writing process. Drafts, notes, sources you read, and the file's own history. How to prove you didn't use AI: version history in Word and Google Docs covers exporting it. Being able to talk through how you built the argument is something a chatbot's verdict cannot see.
  4. If you cite this study, cite both ends. "Chatbots are wrong a third of the time" is the previous generation only. The newest models in this test labelled 0.1% to 8.6% of human texts as machine-written, the previous generation far more. And a rate describes a large pile of texts, not your one essay, which is the point of What a 1% false positive rate means when a university submits 75,000 papers.
  5. Don't answer with another screenshot. A different chatbot saying "human" carries every limit above. Your drafts are evidence of how the essay was written; a second chatbot opinion is not.

Where HumanPen fits

A chatbot's answer is not a report and has no flagged passages in it, and reworking an essay that is already being questioned would not settle anything. HumanPen's humanize option is built for a different case: a draft you have not handed in yet, a Turnitin or iThenticate report on it, and a course that permits tool-assisted rewording. Check that last part in your course rules first. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file.

Frequently asked questions

Does this mean ChatGPT can spot AI writing now? Not on this evidence alone. The paper describes testing through APIs, not the ChatGPT app: GPT-5.5 and 14 other models, each given one fixed question. In that setup the newest models wrongly labelled 3.6% of human texts on average and missed 6.1% to 13.4% of AI texts, depending on which generation wrote them.

Is a newer chatbot always more reliable? Not necessarily. The previous generation labelled more human texts "machine" than the oldest one did (36.5% against 13.0%). And one previous-generation model, the Qwen 72B, did so for 4.4% of human texts, less than three of the newest models (Claude Opus 4.7 at 6.1%, MiMo-V2.5-Pro at 7.9%, DeepSeek-V4-Pro at 8.6%), though it got there by answering "human" almost every time.

Has this study been peer reviewed? No. It was posted to arXiv as a preprint on 28 September 2026. The authors have released the texts, the prompts, every model's raw answers and their analysis code.

Sources

  • Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Ruizhe Li and Shujun Li, Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis, arXiv:2609.34691v1, submitted 28 September 2026 (preprint, not peer reviewed); read 2 October 2026 (Tables 1 and 5; Sections 3, 4 and 5; Appendix A.2 and A.3).
  • The authors' released data and results, hyyuan/detect-llm-generated-texts on GitHub: per-model rates from `dataset/results/eval/paper_reproducible/detector_error_profiles.csv`, human texts from `dataset/human-generated-text/human.json`; read 2 October 2026.

KEEP READING