AI Detector Minimum Word Count: What Five Vendors Publish

Pasting the one paragraph you just rewrote into a free checker is the most natural thing to do, and it is the case every one of these vendors warns about in its own documentation. Five of them publish a minimum length. One of them published two experiments in which its own staff wrote a short passage by hand and the tool called it AI.

HumanPen Team

· 10 min read

How much text does an AI detector need before its number means anything?

Every detector here publishes a floor, and a single paragraph sits under most of them. These are the vendors' own figures, taken from their own documentation and checked on 18 September 2026:

ToolWhat it publishes
Copyleaks — hard minimum"For the AI Detector to get an accurate reading, it requires a minimum of 350 characters" (characters, not words)
Copyleaks — suggested sample"we suggest testing text containing an average of 350 words"
Winston AI"Minimum 300 characters. Texts under 600 characters may produce unreliable results and should be avoided."
Originality.ai100-word minimum on the website; in the API there is no minimum, but "accuracy is decreased for texts below 100 words"
ZeroGPT"At least 150–200 words; longer text (500–1,000+) improves stability."
TurnitinNo AI score at all under 300 words of prose

Two of those floors are written in characters and the rest in words, which is why they are easy to mix up. The conversion is quick: in English prose a word runs about six characters once you count the space after it — across the 5,357 paragraphs published on this site the median is 6.0. Converted, every floor on the table becomes a word count: Winston's 300-character minimum is about 50 words and its avoid-below-600-characters line about 100; Copyleaks' 350-character minimum is about 58, and the sample it suggests is 350 words; Originality's is 100; ZeroGPT asks for at least 150 to 200, and says 500 to 1,000-plus is steadier still; Turnitin's is 300. Hold your own passage against those. For scale, we counted the same 5,357 paragraphs against the same floors: 54% reach Winston's 300 characters, 39% reach Copyleaks' 350, 6% reach Originality's 100 words, 0.5% reach ZeroGPT's 150, and none of them reaches the 350-word sample Copyleaks suggests. The median paragraph is 51 words and 312 characters — past one of the seven floors on the table, Winston's 300 characters, and short of the other six.

Worth separating two things that get run together. A floor is the point below which the vendor stops standing behind its own number — not, in every case, the point below which nothing comes back. Some of these tools return a percentage anyway: Originality.ai says scans in its API "have no word count minimum" and then tells you accuracy drops below 100 words. Others simply refuse, and Turnitin is one of them. So a number on screen is not by itself evidence that the tool considered the sample large enough.

Why the paragraph you just rewrote is the hardest case

Fix the passage that was highlighted, paste that passage into a checker, read the number, decide whether you are done. Originality.ai's help centre puts this exact move near the top of its list of false-positive causes:

You are more likely to receive false positives by only scanning part of the entire piece of content: a single paragraph, the intro, the conclusion, half of the paper... etc.

The next sentence says what to do instead: "We can fix a lot of false positives by scanning the entire piece of content instead of a snippet. This just gives our tool more context and more content to go off of."

The same page carries two experiments the vendor ran on itself, which is rarer and more useful than the advice:

  • A 50-word blog introduction, written by hand by the person writing the help article, came back 79% AI.
  • A 107-word conclusion, also written by hand, came back 69% AI. The page's explanation is that a conclusion is formulaic as well as short: "You summarize everything before, give the reader more options and places to click, and finish up with any final calls to action."

Both numbers came from human-written text. Neither is evidence about your writing. They are evidence about what a short, structurally predictable extract does to a score, and the introduction and conclusion of an essay are both short and structurally predictable by design.

Short extracts also change what the formatting of your text is worth. Copyleaks' limitations page says its detector reads thousands of patterns, "not just words or formatting", and then adds the part that matters for a short sample: "Numbering, bullets, and outline-style headers are only one small part of the overall signal, but in very short texts they can have a larger relative impact to overall AI detection scores." A methods paragraph with three numbered steps is exactly that shape.

There is a second reason the rewritten paragraph is the worst possible sample, and it has nothing to do with length. It is the part of your document that has just been edited most heavily. Originality's page is blunt about what tool use does to its score: "if a tool interacted with your content in any way, it will raise the AI score", and it names grammar checkers alongside chatbots. Whatever you think of that as a policy, it means the paragraph you have been working on is the paragraph carrying the most editing signal, and you are asking the tool to judge it alone.

A sentence-level number is weaker than the document number

Most detectors now colour individual sentences, which invites reading the report sentence by sentence. Two vendors say directly that those smaller readings carry less weight than the overall one.

Winston AI's API documentation, describing the array of per-sentence scores it returns, adds: "Please note that assessments on smaller samples are less accurate than the general score."

GPTZero's guidance on its sentence highlighting goes further, and it answers a question people ask us constantly — why editing the highlighted sentences did not move the score:

Some sentences in an otherwise AI or human document might not show up as highlighted. They may still have an effect on your score, but they aren't more likely than other parts of your document.

And then, with the reason attached: "You may find that adjusting an entire document changes your score more than just adjusting the sentences. This is because our detector holistically analyzes the document, and its prediction can depend on a pattern that affects the bulk of the text."

That is the vendor saying the highlights are not a complete map of what produced the number. What AI detectors actually measure covers how these scores are produced, and why the familiar explanation of them is out of date; the practical point here is narrower. If you treat the highlighted sentences as the full list of what needs attention, you are working from a partial map, by the vendor's own description.

The same text, the same tool, a different answer

Length is not the only thing that moves a reading. Copyleaks lets an institution choose one of three sensitivity settings, and publishes what each does to its own error rates:

SettingHow Copyleaks describes itFalse positivesFalse negatives
Extra Safe"Designed to minimize false positives by using additional AI detection based filters."0.010%0.5765%
Balanced (default)"Ideal for detecting AI content while minimizing false positives."0.0295%0.174%
Extra Sensitive"Our most sensitive model, designed to flag AI text that was put through a 'humanizer' or text spinner."0.1973%0.063%

Read the first and last rows together. Moving from the most conservative setting to the most sensitive multiplies the published false-positive rate by about twenty, and it is a switch on someone else's screen. You cannot see which position it is in, and nothing in a report you are handed will tell you.

So when your own check and your school's check disagree, there are at least three ordinary explanations before anyone gets to "one of them is broken": a different tool, a different setting on the same tool, and a different amount of text. Why two AI detectors give different scores goes through the rest.

The reading that counts is on the whole file

Whatever you run at home, the number with consequences is produced by your institution's tool, on the file you submit, at its settings. Turnitin will not produce an AI score at all for a submission under 300 words of prose, and on documents just above that line it describes its own prediction as close to all-or-nothing — what that does to short assignments is its own subject.

Turnitin's newer Clarity assignments publish a working range in the same spirit: "For Turnitin Clarity to work optimally, we recommend a minimum word count of 300 words and a maximum of approximately 2,500 words", and where AI writing detection is switched on, "submissions must be between 300 and 30,000 words."

Being able to run your own check at all depends on what your institution has enabled and on which screens you can reach, which is a separate question with its own answer: how to check your work before you submit it.

What to do instead

  1. Scan the whole document, not the passage. This is the single change that matches every vendor's own advice: no new tool, no new account, no new setting, just the same check on more text. On tools that meter by length it does cost more — Winston's API documentation says "Each word that is processed by the API consumes one credit" — and it is still the reading their own pages tell you to take.
  2. Compare like with like. A before-and-after comparison only means something if the tool, the settings and the amount of text are the same on both sides. Changing the sample size between the two readings changes the reading on its own.
  3. Expect the number to move without you touching the text. Models are re-versioned. Winston's API documentation lists twenty-two selectable versions, from 2.0 to 4.18, and defaults to the newest. A score from three weeks ago and a score from today may not come from the same model.
  4. Do not treat a number from your own check as a verdict, in either direction. A low reading at home is not a clearance, because your school's tool is a different tool. A high reading on a 60-word extract is not a finding either: 60 words clears one of the seven floors above and none of the rest.
  5. Keep the report you were actually given. If a flag has already happened, the document that matters is the one your institution produced, not a reading you generated afterwards. Rechecking after a rewrite covers what that comparison can and cannot show.

Frequently asked questions

Is there a word count at which detectors become reliable? None of these vendors claims a point at which its output becomes certain. What they publish is a floor below which they say it is unreliable, and a direction: longer is steadier. Copyleaks puts it as "the accuracy of our detection increases as the text length increases"; ZeroGPT's phrasing is that longer text "improves stability".

My 150-word paragraph came back at 80%. Does that mean it is 80% AI? No, and the confusion is worth clearing up because it is built into how these scores are named. Originality.ai's help centre states it plainly: "an AI score of 90% doesn't mean that AI produced 90% of the content. It means that our tool is 90% confident that AI was used in some part to produce the content". A confidence figure and a proportion-of-the-document figure look identical on screen and mean different things. Turnitin's percentage is the third variant again — a share of qualifying text.

Can I just merge a few paragraphs to get over the floor? ZeroGPT suggests exactly that: "consider merging sections before checking." It gets you over a length threshold, but the reading then describes the merged block, not the paragraph you were worried about. If your question is about one passage, no arrangement of the text answers it cleanly, which is the honest answer rather than a satisfying one.

Do these floors apply to my school's Turnitin report too? The floors above belong to each vendor's own detector. Turnitin has its own: no score under 300 words of prose, and only prose sentences count towards the percentage at all. Lists, tables and bullet points are excluded from what it measures, which is why the highlights and the number can look out of step.

Sources

KEEP READING