英文论文 AI 率太高怎么办?先搞懂报告在说什么

When Turnitin says your English essay has a high AI rate, the first step is not to lower it. It is to find out what the number actually is, which display rule produced it, and which side of the model's probability split your word choices landed on. The fix is not a dial called perplexity. The fix is moving your word probability back toward your own human pattern.

HumanPen Team

· 35 min read

The short answer

Turnitin's AI writing report shows no number at all between 1% and 19%; it displays an asterisk and withholds the highlights. Only at 20% and above does a precise percentage appear, with highlighted passages you can actually look at. So when you say "my AI rate is too high", the first thing to do is not to lower it. It is to find out whether you are looking at a number or a star, because only the number comes with something you can act on.

The second step is to ask why it is high. Turnitin's model cuts your text into overlapping segments, gives each one a probability between 0 and 1, pools those scores, and aggregates them into a document percentage. If your word choices fall on the AI-generated side of the model's probability split, your score goes up. Writing in English as a Chinese native speaker has three statistical tendencies that push your probability toward the AI side. We break each one open below.

The third step is revision. The direction is not to lower a metric called perplexity, because Turnitin does not explicitly compute one. The direction is to move your word probability back toward your own human pattern: use the words you actually reach for, use sentence structures you do not normally produce, and revise the flagged passages one by one rather than swapping synonyms.

What "too high" actually means

Turnitin's AI writing report has a special display rule for low scores. The detection capabilities FAQ puts it this way:

"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."

So if the "AI rate" you are holding is a number with a percent sign, it is at least 20. If you are looking at an asterisk, you are somewhere between 1 and 19, and the report does not tell you where. Zero is the only exact figure displayed below twenty.

This distinction looks pedantic but it decides what you can do next. An asterisk means no highlighted passages, no way to locate what triggered the flag, and nothing to revise against. A number at 20 or above comes with highlights, and highlights are the only thing that let you locate and change the flagged text. So "my AI rate is too high, what do I do" has an actionable answer only when your report shows 20% or above and you can see the highlights. We worked through the report itself in how to read a Turnitin AI writing report.

One more thing to confirm first. Turnitin states that "only instructors and administrators are able to see the indicator", and that "The AI writing detection indicator and report are not visible to students. However, with the PDF download feature, instructors can download and share the AI report with students." If you have a report in your hands, it is because someone decided to download it and give it to you. Confirm you have the full PDF, not a screenshot.

Why it is high: your word probability is on the AI side

Turnitin describes the calculation chain in its FAQ:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

This is a probability chain. Each of your sentences gets a score between 0 and 1 from the model; the closer to 1, the more AI-like. Those scores aggregate into the percentage you see.

So the mechanical meaning of "high AI rate" is specific: your writing, in the model's probability space, falls on the AI-generated side. Not because you used AI, but because the statistical features of your word choices and sentence patterns happen to be closer to the average pattern of AI-generated text.

Here a widespread claim needs correcting. Many answers describe Turnitin's mechanism as "measuring perplexity and burstiness" and then tell you how to lower those two values. Turnitin explicitly denies this in its FAQ:

"Our model is not explicitly programmed to evaluate specific signals such as "burstiness," "perplexity," or other individual metrics sometimes referenced in public discussions."

The sentence directly after it:

"Instead, it learns statistical patterns from our training data."

But this does not mean Turnitin ignores word probability. The same FAQ page says elsewhere:

"Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers."

Those two statements have to be read together. Turnitin does not explicitly compute a score called perplexity, but its classifier is making a judgment at the level of word probability, because perplexity is itself a measure of word probability. The correct statement is: you cannot tune a dial called perplexity, because that dial does not exist as an explicit input. What you can change is your own word probability distribution, so that the model places it back on the human-written side.

Why Chinese native speakers writing English get flagged more

This section is about three statistical tendencies. Each one pushes your word probability toward the AI side. They are not flaws in Chinese native speakers. They are patterns that show up when you are writing in a second language while also managing grammar, vocabulary and argument at the same time.

First, the tendency to reach for "safe" high-frequency words. When writing in English as a second language, the natural instinct is to choose words with stable meanings and low risk of error. Those words are high-frequency words. They are also the words AI-generated text reaches for most often, because language models select by probability and high-frequency words have high probability. The more your word distribution concentrates at the top of the frequency table, the closer it sits to the average AI-generated distribution. The line we already quoted, "Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers", is exactly about this: the classifier reads word probability sequences, and human writers have their own characteristic sequences. High-frequency concentration is not one of them.

Second, low structural variation. Turnitin itself lists the features that tend to produce false positives:

"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."

The next sentence:

"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."

When you are writing in English as a second language, the cognitive budget left for sentence structure is narrow. The result is structural consistency: subject-verb-object, subject-verb-object, subject-verb-object. Low structural variation is exactly the first item on Turnitin's own list of false-positive-prone features. The official advice to educators is to discount high scores on such text, and that sentence has to travel with the first one.

Third, possible use of machine translation or Grammarly. Machine translation output is itself produced by a language model, so its word probability distribution naturally sits on the AI side. Grammarly's situation is more specific. Turnitin's FAQ has a dedicated statement: spelling, grammar and punctuation changes are not flagged "in most cases", but Grammarly's generative features, including draft generation, paraphrasing and summarizing, are not covered by that exemption. The official wording is that content produced using those features "will likely be flagged as AI-generated". So if you used Grammarly's sentence rewrite, those passages being flagged is expected, not surprising.

Why AI detectors are biased against non-native English writers places this family of questions in a wider frame.

How to lower it: change word probability, not perplexity

First, what not to do.

Do not try to lower perplexity. We already quoted Turnitin saying its model is not explicitly programmed to evaluate perplexity as an individual metric. Any advice that tells you to "lower perplexity" starts from a false premise. You cannot lower a value that Turnitin does not explicitly compute. A lot of humanizer marketing copy is built on that false premise.

Do not swap synonyms. Swapping synonyms changes a word, not a word probability distribution. If you replace "important" with "crucial", you have changed one token, but "crucial" is also a high-frequency word in AI-generated text. Your overall distribution has not moved. The model sees the same distribution, just with one slot relabelled. Synonym replacement acts on the surface. The model reads statistics.

Do not rewrite the whole document. Short documents have a documented extreme behaviour:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

The next sentence:

"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

So a half-rewritten short document is especially risky. You rewrote half, the other half gets called AI, and the whole thing comes back high.

Now, what to do.

Get the report first. We already quoted this: students cannot see the indicator, but instructors can download the PDF and share it. You have the right to ask for that PDF. Once you have it, find the highlighted passages. Only reports at 20% and above have highlights. Knowing which passages are flagged is the precondition for revising them.

Revise the flagged passages one by one. The direction is to change sentence structure, not to change words. Take the subject-verb-object pattern you default to and replace it with a structure you do not normally write: a subordinate clause, an inversion, a parenthetical, a cross-sentence reference. These structures change your word probability sequence, which is what the model reads. A sentence's score is also influenced by its neighbours, because segments overlap, so revising one passage shifts the scores of the sentences around it.

Use the words you actually reach for. Resist the pull toward safe high-frequency vocabulary. Use the word you actually want, even if it is less common. The distinctive part of a human writer's word probability sequence is often in the choices that deviate from the top of the frequency table.

Mind the boundary of qualifying text. Turnitin analyses only standard grammatical sentences:

"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."

The next sentence:

"This percentage is not necessarily the percentage of the entire submission."

So the number you see is a percentage over qualifying text, not over the whole file. When you revise, aim at the qualifying text. Editing a bullet list does not move that numerator.

The design trade-off: the model prefers misses over false alarms

One more thing, because it affects how you read your own score. Turnitin discloses the design trade-off in its FAQ:

"In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document. We're comfortable with that since we do not want to incorrectly highlight human-written text as AI-written. For example, if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."

The direction is explicit: the model would rather miss AI text than flag human text. That means the report has a built-in tendency to under-report, not over-report. If you are looking at a high score, that score is on the conservative side of the model's design, not the aggressive side. This does not prove the score is correct. It does mean you should not treat it as a ceiling.

What "changing structure" looks like in practice

"Change the structure" sounds abstract. An example makes it concrete.

  • Original: "The results show that the method is effective."
  • Synonym swap: "The findings demonstrate that the approach is successful." This changes words. The structure is identical. The model's judgment will probably not move.
  • Structural change: "What the results show is not that the method works in general, but that it works under the specific condition we tested." The sentence moves from simple subject-verb-object to an emphasised construction with a turn. The word probability sequence is different. The model sees a different distribution.

Changing structure is not about writing fancy sentences. It is about writing the sentence structures you do not normally produce, because the structures you do not normally produce are exactly where your human-writer probability signature lives.

Five minutes with the report

  1. Confirm you have the PDF, not a screenshot. A screenshot has no date, no highlight positions, and no way to tell which report it is.
  2. Check whether the figure is a percentage or an asterisk. A percentage means 20% or above, and highlights to work from. An asterisk means 1% to 19%, no highlights, and nothing to locate.
  3. Find the specific highlighted passages. Those are the ones to revise. Leave the rest alone. Changing text that was not flagged is wasted effort.
  4. Read each flagged passage and diagnose which tendency produced it. Uniform sentence structure? High-frequency vocabulary? A Grammarly rewrite? Different causes need different fixes.
  5. Change structure, not words. Rewrite the flagged sentences in structures you do not normally use.
  6. Re-check after revising. Detectors change, and Turnitin says its model iterates, so no fixed outcome can be promised. What you can do is run another report after revision and see whether the highlights have receded. Why different AI detectors disagree explains why results from different detectors cannot validate each other.

Frequently asked questions

Does a high AI rate mean I used AI? No. Turnitin's model reads word probability distributions, not usage logs. You can be flagged without using AI, if your vocabulary is high-frequency, your sentence structure is uniform, or you used machine translation. Turnitin itself lists the features that tend to produce false positives, and advises educators to discount high scores on such text.

Can I check my own AI rate? Not through Turnitin's indicator. The official statement is that the indicator and report are not visible to students, and that instructors can download the report as a PDF and share it. Every report you see, you see because someone chose to show you.

Will swapping synonyms lower my AI rate? Probably not. Swapping synonyms changes individual words without changing the word probability distribution. The model reads the distribution, not the individual words. Changing structure is more effective than changing vocabulary.

Will I be flagged for using Grammarly? Spelling, grammar and punctuation changes are not flagged in most cases. But Grammarly's generative features, including draft generation, paraphrasing and summarizing, will likely be flagged as AI-generated. Distinguish between the two.

Are short papers more likely to be flagged? Documents of a few hundred words have predictions that are mostly all-or-nothing, and mixed content can be flagged entirely as AI. Be especially careful with partial rewrites on short documents.

Can a humanizer tool lower my score? This article makes no efficacy claim. Turnitin's design direction is to prefer misses over false alarms, the model iterates, and any promise about a future detection result is not one a tool can honestly make. What we can say is that structural revision is more effective than synonym replacement, and that revising with the report in hand is more effective than revising blind.

KEEP READING