AIGC Detection for English Papers: What It Means and How It Works

"AIGC detection" is the term many Chinese researchers use when asking about AI writing detection for English papers. The detection mechanism is the same regardless of what you call it. Turnitin's FAQ describes a pipeline that classifies overlapping text segments by word probability. Understanding the mechanism helps you read your report accurately.

HumanPen Team

· 10 min read

The Short Answer

AIGC detection for English papers refers to the same process as AI writing detection. Turnitin's detector splits your text into overlapping segments, classifies each segment by the probability that it was human or AI-generated, and aggregates those scores into a document-level percentage. The term "AIGC" (AI-generated content) is commonly used in Chinese academic contexts, but the detection mechanism is the same regardless of what you call it. The detector does not look for specific AI tool signatures or watermarks. It reads statistical patterns in word choice, specifically word probability patterns learned during training. The AI score is completely independent of the similarity (plagiarism) score, and only instructors or administrators can see the AI indicator.

How AIGC Detection Works

Turnitin's FAQ describes the detection pipeline:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

The detector is not scanning for specific phrases that "look like AI." It is classifying text segments using a model that learned statistical patterns from training data. Each segment gets a probability, and those probabilities are pooled and aggregated into a single percentage. The percentage represents the proportion of qualifying text that the model classifies as likely AI-generated.

What the Model Reads

Two statements from the FAQ clarify the model's approach. First: "Our model is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions." The next sentence: "Instead, it learns statistical patterns from our training data." Second: "Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers."

The model does not compute named metrics like burstiness or perplexity. But it is trained on word probability data. The patterns it learns are expressed in how words are chosen and sequenced, not in surface features like sentence length variation or vocabulary diversity. This is why surface-level edits, like changing a few words or adding transitions, may not change the score. The detector reads deeper statistical features of the word sequences, which is the subject of what AI detectors measure beyond perplexity and burstiness.

Only Prose Is Analyzed

The FAQ specifies what counts as analyzable text:

"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."

The next sentence: "This percentage is not necessarily the percentage of the entire submission."

For English academic papers, this means that your methodology tables, your equation blocks, your figure captions, and your reference list are largely excluded from AI analysis. The percentage applies only to the prose portions of your manuscript. A paper with extensive tables and equations will have less qualifying text, which means the AI percentage represents a smaller portion of the document than it might appear. How to read a Turnitin AI writing report works through where that gap shows up on the page.

AIGC Score vs Similarity Score

The two scores are separate measurements:

"The Similarity score and the AI writing detection percentage are completely independent and do not influence each other."

Your similarity score reflects how much of your text matches existing sources in the database. Your AIGC score reflects how much of your text the model classifies as likely AI-generated. A paper can have high similarity (from properly cited sources) and a low AIGC score, or low similarity and a high AIGC score. Working on one score does not affect the other. If you are trying to lower your AIGC score, lowering your similarity score will not help, and vice versa; reading the two scores as the same thing is where this usually goes wrong.

What to Do If Your AIGC Score Is High

To summarize what we have covered:

  • AIGC detection for English papers works the same way as AI writing detection: classifying word probability patterns in prose segments.
  • The model does not compute burstiness or perplexity as named metrics, but it is trained on word probability data.
  • Only prose sentences are analyzed. Tables, equations, and reference lists are largely excluded.
  • The AIGC score and similarity score are completely independent. Changing one does not affect the other.
  • Only instructors and administrators can see the AI indicator. Students can request the PDF version from their instructor.

If you have a Turnitin or iThenticate report showing which passages were flagged, import the report and work on those specific passages. Eligible passages can be re-run at no charge.

KEEP READING