为什么引言和结论最容易被标记为 AI

很多同学注意到引言和结论是最常被标记为 AI 的部分。这不是巧合。Turnitin 2023 年 5 月的 release notes 描述了文档首尾几句误报率更高的现象。检测逻辑已更新以减少这种情况,但泛化引言和结论写作的模式仍然存在。

HumanPen 团队

· 10 分钟

简短回答

引言和结论比正文段落更容易被标记为 AI。Turnitin 2023 年 5 月的 release notes 声明:"Since launch, we have observed a higher incidence of false positive detection in the first few or last few sentences of a document. Many times these sentences consist of introduction or conclusion content written in a generic way." 检测逻辑被更新以减少这些误报,但泛化、公式化的引言和结论写作模式仍然存在。如果你的引言以宽泛的陈述开始、结论以可预测的语言重申论点,这些段落可能产出的词概率特征和 AI 生成文本有重叠,即使在 2023 年改进之后。

Release notes 说了什么

2023 年 5 月的 release note 描述了现象和回应:

"Since launch, we have observed a higher incidence of false positive detection in the first few or last few sentences of a document. Many times these sentences consist of introduction or conclusion content written in a generic way. As a result, we have changed our detection logic to help reduce these false positives."

下一句:"We also worked on making our segment boundaries detection more precise which could lead in some rare cases to change of boundaries compared with a previous version."

这是 2023 年的一项改进,不是对现行缺陷的描述。检测逻辑被更改以帮助减少这些位置的误报。但这条记录告诉我们一件重要的事:文档的首尾几句本身更容易误报。改进减少了问题。它没有消除问题。

为什么引言和结论共享 AI 模式

Release note 指向"introduction or conclusion content written in a generic way." 泛化内容是这里的连接组织。引言经常以宽泛背景开始:"In recent years, X has become an important topic." 结论经常重申论点:"This paper has shown that X leads to Y." 这些构造成分是可预测的、公式化的、结构上统一的。Turnitin 的 FAQ 把"content without a lot of structural variation"列为误判易发特征。泛化的引言和结论恰好是这类内容。这些章节中的词概率模式往往比正文中更可预测,因为正文论证更具体、语言更多变。这件事的一般形式是为什么写得好的文章反而被判成 AI

检测器怎么处理文档边界

检测流水线通过切片文本工作:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score."

Release note 还提到"segment boundaries detection"。文档的首尾几句位于片段结构的边缘。在 2023 年改进之前,这些位置误报率更高。改进让片段边界检测更精确了,这有帮助。但如果内容本身写得泛化,它仍然可能产出触发分类为 AI 式的词概率模式。那些模式是什么,见Turnitin 到底用不用 perplexity 和 burstiness 来判定 AI

短文档放大这个问题

Turnitin 的 FAQ 还指出:

"In shorter documents where there are only a few hundred words, the prediction will be mostly "all or nothing" because we're predicting on a single segment without the opportunity to overlap."

下一句:"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

在短论文中,引言和结论可能占总文本的很大比例。如果两者都被标记,短文档的全有或全无行为意味着整篇文档可能被分类为 AI 生成。这是一个极端结果,但它有助于解释为什么一篇有泛化引言和结论的短论文可能拿到出乎意料的高分。Turnitin 报告显示 100% AI,是怎么来的把这个极端情况从头走了一遍。

怎么办

总结一下我们讲的内容:

  • Turnitin 2023 年 5 月的 release notes 描述了文档首尾几句误报率更高的现象。
  • 检测逻辑在 2023 年更新以减少这种情况,但泛化引言和结论写作的底层模式仍然存在。
  • 泛化的引言和结论共享 FAQ 列为误判易发的"低结构变化"特征。
  • 短文档放大问题因为引言和结论占总文本比例更大。
  • 让引言和结论更具体、更少公式化、结构更多变可以帮助应对这个模式。

如果你有 Turnitin 报告显示哪些段落被标记了,可以导入报告专门处理那些段落。符合条件时可以免费继续降 AI。

继续阅读