GPTZero vs Turnitin:为什么分数对不上,以及每个工具实际在测什么
如果你把同一段文本分别提交给 GPTZero 和 Turnitin,分数往往对不上。这不代表哪个工具坏了。不同的检测器使用不同的模型、不同的训练数据、不同的评分规则。以下是 Turnitin 文档关于其检测机制的实际说明,以及为什么跨工具比较不可靠。
HumanPen 团队
· 14 分钟
不同检测器使用不同模型
GPTZero 和 Turnitin 是由不同团队构建的独立系统。各自有自己的底层模型、自己的训练数据、自己的文本分类方式。当两个检测器使用不同模型时,它们对同一文档的输出不会一致。GPTZero 打了 60% AI 分的文本,Turnitin 可能给 30%,反过来也一样。两个分数都不一定错。它们是用不同模型在不同数据上训练后给出的概率估计。
这是一个根本性的问题:不存在一个所有工具都会趋同的通用 AI 检测分数,为什么同一段文字在不同检测器上分数差很多讲的就是这个问题的形状。每个工具基于自己模型的判断给出自己的估计。把不同工具的分数当作在测同一个东西、用同一种单位来比较,会导致混乱。
我们不声称哪个工具更准确。我们只指出文档描述了 Turnitin 检测器的运作方式,而这个机制与其他工具不同。
Turnitin 检测器如何给文档评分
Turnitin 的文档描述了一个具体的多步骤流程,用于得出文档级别的 AI 写作分数:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."
分数是对段落级别预测的汇总。段落之间有重叠,句子分数在最终汇总前会被合并。这是一种特定的方法论。其他检测器可能用不同方式分段、使用不同预测单元,或用不同方式汇总分数。
Turnitin 文档还澄清了一件事:AI 写作检测百分比和相似度分数是完全独立的测量。"The Similarity score and the AI writing detection percentage are completely independent and do not influence each other." 相似度分数表示提交文档与 Turnitin 供相似度检查的内容库比对后匹配到的文本比例。高相似度不代表高 AI 分,高 AI 分也不代表高相似度。Turnitin AI 写作报告和相似度报告有什么区别把两者摆在一起对照过。
Turnitin 的误报率目标
Turnitin 在文档中给出了具体的误报率目标:
"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."
文档随后做了改述:"In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents." 注意这个改述丢掉了"AI 写作超过 20%"这个限定条件。带 20% 阈值的版本是更精确的表述。
为了验证这个比率,Turnitin 描述了一个测试流程:"To bolster our testing framework and diagnose statistical trends of false positives, before every update or new model release, we perform tests on over 700,000 additional academic papers that were written before the release of ChatGPT to further validate our less than 1% false positive rate."
这是 Turnitin 声明的目标和验证流程。放到一所大学一年的提交量上,这 1% 具体是多少人,见1% 的误判率意味着什么。其他检测器可能有不同的目标、不同的验证方法,或者可能没有公布可比的数字。我们无法基于这些信息在工具之间做直接的准确率比较。我们只能报告 Turnitin 文档关于其自身系统的说明。
短文档的表现不同
文档长度会影响 Turnitin 检测器的行为。对于短文本,评分机制会简化并可能产生极端结果:
"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."
"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."
原因是结构性的。在较长文档中平滑分数的重叠机制在只有一段可预测时不起作用。几百字的混合人类和 AI 文本可能被判为 100% AI,这是Turnitin 报告显示 100% AI,是怎么来的里列出的路径之一。这是评分方法在短输入上的局限,不是对文本成分的确定性判断。
如果 GPTZero 对一篇短文档给出了与 Turnitin 不同的分数,短文档的全有或全无行为是一个可能的解释。两个工具处理短文本的方式不同,因为底层评分架构不同。
1-19% 的星号区间
Turnitin 对于 1% 到 19% 之间的 AI 检测结果不显示数字分数。文档解释:
"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."
这意味着如果 Turnitin 的模型估算 AI 内容为比如 12%,报告会显示星号而不是数字。文本上也不会有高亮。这样做的目的是避免在低置信度区间呈现可能的误报,更细的说明见Turnitin AI 率上的星号(*%)是什么意思。
GPTZero 和其他工具在这个区间可能以不同方式显示分数。有些工具对任何非零结果都显示完整百分比。Turnitin 则在 20% 以下刻意隐藏数字。这也是分数跨工具对不上的另一个原因:一个工具显示数字分数的地方,另一个工具可能只显示一个星号。
这对你意味着什么
如果你在同一篇文档上比较 GPTZero 和 Turnitin 的分数,请注意以下几点:
- 不同模型,不同分数。 GPTZero 和 Turnitin 使用各自的检测模型,训练数据不同,分数不期望一致。
- Turnitin 通过分段汇总评分。 文本先分段、逐段评分、再汇总。AI 分数和相似度分数完全独立。
- Turnitin 目标误报率低于 1%(针对 AI 写作超过 20% 的文档),通过 70 万篇 ChatGPT 发布前撰写的学术论文验证。
- 短文档得到全有或全无的分数。 几百字可能产生 0% 或 100% 的结果,即使文本是混合的。
- 20% 以下的分数显示为星号。 在 1-19% 区间不显示数字百分比或高亮。
符合条件时可以免费继续降 AI。
继续阅读