Undetectable AI vs Turnitin:文档实际怎么说的

如果你把同一段文本分别提交给 Undetectable AI 的检测器和 Turnitin,分数往往不同。这是正常的。两个工具使用不同的模型、不同的训练数据、不同的评分规则。以下是 Turnitin 文档关于其检测器运作方式的说法,以及为什么两个分数不会一致。

HumanPen 团队

· 14 分钟

Undetectable AI 和 Turnitin 是独立系统

Undetectable AI 的检测器和 Turnitin 的 AI 写作检测由不同团队构建,使用不同的底层模型。每个系统在自己的数据上训练,用自己的方式分类文本。当两个检测器使用不同模型时,它们对同一段文本会产生不同分数。Undetectable AI 判为 70% AI 的段落,Turnitin 可能给 25%,反过来也一样。两个结果都不一定错。它们是不同模型在不同数据上训练后给出的概率估计。

不存在一个所有工具都会趋同的通用 AI 检测分数,这一点在为什么同一段文字在不同检测器上分数差很多里展开过。每个检测器给出自己的估计。把不同工具的分数当作在测同一个东西、用同一种单位来比较,会导致混乱。

我们不声称哪个工具更准确。我们只描述 Turnitin 文档关于其自身系统的说法,以及为什么该系统产生的分数不会与 Undetectable AI 的一致。

Turnitin 的评分流程怎么运作

Turnitin 文档描述了得出 AI 写作分数的具体流程:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

流程包括提取句子、分段为重叠单元、逐段评分、合并句子级分数、汇总为文档级百分比。这是一种特定的架构。其他检测器,包括 Undetectable AI,可能使用不同的分段策略、不同的预测单元或不同的汇总方法。结果是两个工具对同一段文本不会给出一致的数字。

Turnitin 文档还澄清,AI 写作检测百分比和相似度分数是分开的。"The Similarity score and the AI writing detection percentage are completely independent and do not influence each other." 相似度分数表示提交文档与 Turnitin 供相似度检查的内容库比对后匹配到的文本比例。一篇文档可以高相似度低 AI 分,也可以高 AI 分低相似度。Turnitin AI 写作报告和相似度报告有什么区别说明了两个数字各自在数什么。这两个数字互不影响。

Turnitin 公布的误报率

Turnitin 公布了具体的误报率目标:

"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."

文档做了改述:"In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents." 这个改述丢掉了原说法中"AI 写作超过 20%"的限定条件。带 20% 阈值的版本更精确。

为了验证这个目标,Turnitin 描述了其测试框架:"To bolster our testing framework and diagnose statistical trends of false positives, before every update or new model release, we perform tests on over 700,000 additional academic papers that were written before the release of ChatGPT to further validate our less than 1% false positive rate."

这是 Turnitin 关于自身系统的说法。这 1% 摊到一年的提交量上是多少份,见1% 的误判率意味着什么。Undetectable AI 可能有不同的准确率目标、不同的验证流程,或者可能没有公布可比的数字。我们无法在两个工具之间做直接的准确率比较。我们只能报告 Turnitin 文档的陈述。

短文档的行为

文档长度会影响 Turnitin 检测器处理文本的方式。对于短文档,评分机制的行为不同:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

在较长文档中平滑分数的重叠机制在文本只有一段时不起作用。几百字的混合人类和 AI 内容可能返回 100% 的分数,这是Turnitin 报告显示 100% AI,是怎么来的背后的机制之一。这是评分方法在短输入上的结构性行为。

如果 Undetectable AI 对同一篇短文档给出了不同分数,部分原因是两个工具处理短输入的方式不同。

低分如何显示

Turnitin 在 20% 以下隐藏数字分数。文档指出:

"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."

Turnitin 模型估算为 8% AI 的文档,报告里会显示星号,而不是数字 8%。文本上也不会有高亮。这样做的理由是避免在低置信度区间呈现可能的误报,更详细的说明见Turnitin AI 率上的星号(*%)是什么意思

Undetectable AI 在这个区间可能以完整百分比显示分数。其他工具可能对任何非零水平都做高亮。Turnitin 则在 20% 以下刻意隐藏数字和高亮。这种评分约定意味着两个工具对同一篇低 AI 文档会呈现非常不同的输出。

这对你意味着什么

如果你在同一篇文档上比较 Undetectable AI 和 Turnitin 的结果,以下是总结:

  • 不同模型,不同分数。 Undetectable AI 和 Turnitin 使用不同的检测模型,分数不期望一致。
  • Turnitin 通过分段和汇总评分。 句子被提取、分段为重叠单元、评分、合并、汇总。AI 分数和相似度分数完全独立。
  • Turnitin 目标误报率低于 1%(针对 AI 写作超过 20% 的文档),通过 70 万篇 ChatGPT 发布前撰写的学术论文验证。
  • 短文档得到全有或全无的分数。 几百字可能返回 0% 或 100%,即使文本是混合的。
  • 20% 以下的分数显示为星号。 在 1-19% 区间不显示数字百分比或高亮。

符合条件时可以免费继续降 AI。

继续阅读