为什么论文不同段落 AI 分数不一样?Turnitin 检测流水线的一段一段解读

你打开 AI Writing Report,一个段落 60%,另一段显示 0%——或者高亮条盖满第 3 页,第 5 页什么都没有。第一反应是报告坏了。不是。Turnitin 是逐句打分,再把句子分数聚合成分段视图和整体百分比;段落长度和句子长相不同,各段分数自然不同。这篇文章讲官方机制、高亮怎么读,以及为什么 AI 百分比永远不会等于相似度分数。

HumanPen 团队

· 5 分钟

短回答

同一篇论文的不同段落得到不同 AI 分数,是因为 Turnitin 逐句给文本打分,再把句子分数汇总成整体百分比。官方指南写的流水线是:抽取句子、切成重叠分段、给每个分段打分、每句继承分段分数、句子分数再聚合成文档分数。短而密集的段落和长而口语化的段落,在 Submission Breakdown 里的样子天然不同——这是机制按设计在工作,不是报告出故障。

检测器实际怎么读你的论文

官方「Using the AI Writing Report」指南这样描述检测过程:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis."(当论文提交至Turnitin时,系统会从提交的内容中提取句子,并将其分割成重叠的片段以进行预测分析。)
"Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."(每个片段由AI检测模型进行分类,并赋予一个介于0到1之间的值,表示文本可能由人类或AI生成的概率。)

Turnitin 自己的博客补充了粒度细节:

"The submission is first broken into segments of text that are roughly a few hundred words (about five to ten sentences). Those segments are then overlapped with each other to capture each sentence in context."(提交的文档首先被分割成大约几百个单词(约五到十个句子)的文本片段。然后这些片段相互重叠,以便在上下文中捕捉每个句子。)

所以你看到的报告是自下而上建出来的:句子级打分 → 分段级视图 → 顶部单一百分比。

为什么各段会不一样

官方指南继续:

"Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."(这些片段中每个符合条件的句子都会继承该片段的分数。由于片段存在重叠,部分句子可能会有多个分数,这些分数随后会合并为一个单一分数。这些句子分数会被进一步汇总,用于计算整个文档的 AI writing score。)

这个机制带来三个直接推论:

  • 段落按「句子的长相」打分——满是短句、模板化句子的段落,和长句、变化多的段落分数不同。
  • 段落长度有影响:Submission Breakdown 按「该区域内多少 qualifying 文本被标」来涂色。5 页的段落当然比 1 页的段落承载更多高亮,即使逐句模式相近。
  • 重叠 pooling 意味着边界上的句子可能带混合分数——所以高亮的精确边界不是精确的判决边界。

这些都不是报告错误,是 Turnitin 文档里写的聚合方式。

AI 百分比与相似度分数互相独立

最常见的误读是把 AI「分数」和相似度百分比当成同一件事的两次读数。官方指南把两者的分离写得明明白白:

"The percentage generated by Turnitin's AI writing detection model is different from and independent of the similarity score."(Turnitin 的 AI 写作检测模型生成的百分比与 similarity score 不同且相互独立。)
"AI writing highlights are not visible in the Similarity Report."(AI 写作高亮在 Similarity Report 中不可见。)

一个段落相似度 40%、AI 0%,并不矛盾:一个数字数的是匹配文本,另一个是检测模型对句子的判断。文档各段 AI 分数不同的原因,和 AI 分与相似度分不同的原因是一样的——它们在回答不同的问题。

读 Submission Breakdown 和高亮

官方指南描述报告顶部的交互条:

"The overall percentage of text likely detected as AI is detailed in the Submission Breakdown."(可能被检测为 AI 的文本总体百分比在 Submission Breakdown 中有详细说明。)
"The bar is interactive and allows you to select each highlight to bring the corresponding text and page of the submission into focus."(该指示条是交互式的,允许您选择每个高亮部分,以聚焦到提交内容的相应文本和页面。)

读自己报告时的两个实际注意点:

  • 高亮句是蓝色;2026-08-04 起 Turnitin 把 AI-generated 和 AI-paraphrased 合并成单一蓝色类别(旧报告是青/紫两色)。
  • 只有 qualifying prose 参与计算——列表、要点、表格、诗、脚本、代码都不计入。官方原句:"The model does not reliably detect AI-generated text in the form of non-prose, such as poetry, scripts, or code, nor does it detect short-form/unconventional writing such as bullet points, tables, etc."

所以「空白」段落有时不是没文本,而是没有 qualifying 散文。报告对要点和表格什么都没说,因为它从来不算它们。

恐慌之前先知道三种状态

官方指南的指示器状态部分覆盖了不像普通数字的百分比:

"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores above 0% and below the 20% threshold in the report. When this occurs, it is now indicated with an asterisk (%) and no percentage is attributed."(为了避免潜在的假阳性发生率,在报告中,对于高于 0% 且低于 20% 阈值的 AI 检测分数,不分配分数或高亮显示。当发生这种情况时,现在会以星号 (%) 表示,且不分配任何百分比。)
"0% detected as AI … our testing has found that there is a higher incidence of false positives when the percentage is between 0 and 19."(0% 被检测为 AI……我们的测试发现,当百分比在 0 到 19 之间时,假阳性的发生率更高。)

这改变了你读分段差异的方式:显示 0% 的段落,可能是真零,也可能是一个低于 20% 的分数被报告拒绝显示——报告本身无法告诉你到底是哪种。这是设计,不是巧合。

底线

一份显示各段分数不同的 Turnitin AI 报告,是检测流水线在做 Turnitin 文档里写明的事:逐句打分、重叠分段聚合、整体数是聚合结果——不是对整篇文档的统一判决。段落不同是因为句子特征和长度不同。AI 百分比与相似度分数互相独立,高亮只覆盖 qualifying prose,所以要点、表格、代码完全不算。读自己的报告时:把高亮当成调查线索,而不是每个段落的固定标签。

继续阅读