Originality.ai vs Turnitin AI 检测:文档实际说了什么

同学们有时用 Originality.ai 跑论文拿到的分数和 Turnitin 不一样。这是正常的。不同检测器用不同模型训练于不同数据。Turnitin 的 FAQ 描述了它的机制、误报目标和低分的星号规则。以下是文档说了什么。

HumanPen 团队

· 10 分钟

简短回答

Originality.ai 和 Turnitin 是不同的 AI 检测器,基于不同模型、不同训练数据和不同分类阈值。一个的分数不能预测另一个的。Turnitin 的 FAQ 描述了它的机制:文本被切成重叠片段,每个用 0 到 1 之间的概率分分类。FAQ 还声明对 AI 占比超过 20% 的文档误报率目标低于 1%,并对 1-19% 范围的分数使用星号规则。Originality.ai 发布自己的准确率声称,但那些是厂商自报的,无法独立对照 Turnitin 文档验证。重要的分数是你学校使用的工具的分数。如果你的学校用 Turnitin,任何第三方分数都不能告诉你 Turnitin 分数会是多少。

Turnitin 的检测器怎么工作

FAQ 描述检测过程:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."

片段重叠,意味着句子可以接收多个分数并汇总在一起。这种汇总会平滑单个预测错误并贡献到整体文档分。Originality.ai 用自己的分类方法,其在 Turnitin 资料中没有被记录。两个系统构建方式不同,评估文本的方式也不同。

Turnitin 的误报目标

FAQ 声明了一个具体目标:

"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."

下一句:"In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents."

这个目标专属于 Turnitin,适用于 AI 占比超过 20% 的文档。Originality.ai 可能有不同的误报率,但比较两个数字需要理解每个厂商如何定义和衡量误报。我们不就第三方检测器做准确率声称。文档告诉我们的是 Turnitin 对自己的模型明确声明了这个目标。这个目标放到规模上要付多少代价,当一所大学一年交上去 75,000 篇论文,1% 的误判率意味着什么算过。

星号规则

Turnitin 对低分有特定的显示规则:

"To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed."

这意味着显示"%"的 Turnitin 分数表明检测到低于 20% 的 AI,但确切百分比被保留以避免在误报易发范围夸大准确性。Originality.ai 不使用这个规则,会在此范围显示具体百分比。Turnitin 上显示"%"的论文可能在 Originality.ai 上显示"15%"。这不意味着一个工具比另一个更准确。它意味着它们在低置信范围的处理方式不同。Turnitin 这一侧的细节在Turnitin AI 率上的星号(*%)是什么意思

为什么不同工具的分数不一致

不同 AI 检测器在同一文本上产出不同分数,因为它们用不同模型、不同训练数据和不同阈值。Turnitin 的 FAQ 描述了它的模型如何把重叠片段分数汇总成文档级百分比。另一个检测器可能用不同的切分策略、不同的评分尺度或不同的聚合方法。FAQ 还指出短文档产生全有或全无的预测:"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap." Originality.ai 在短文档上可能有也可能没有同样的行为。

AI 分和相似度分在 Turnitin 上也是独立的:"The Similarity score and the AI writing detection percentage are completely independent and do not influence each other." 把 AI 和相似度合成一个分数的检测器产出的数字无法和 Turnitin 分开的 AI 百分比比较。为什么同一段文字在不同检测器上分数差很多把这些原因收在一处。

这对你意味着什么

总结一下我们讲的内容:

  • Originality.ai 和 Turnitin 是不同的检测器,用不同模型。分数不会一致。
  • Turnitin 的检测器把文本切成重叠片段并分配 0 到 1 之间的概率分。
  • Turnitin 的误报率目标对 AI 占比超过 20% 的文档低于 1%。
  • Turnitin 对 1-19% 范围的分数使用星号以避免夸大准确性。
  • 短文档因单片段评估得到全有或全无的预测。
  • 重要的分数是你学校使用的工具的分数。

如果收到 Turnitin AI 报告并想处理被标记的段落,可以导入报告处理。符合条件时可以免费继续降 AI。

继续阅读