Turnitin 能检测 GPT-5 和其他新 AI 模型吗

同学们想知道最新的 AI 模型能不能绕过 Turnitin。FAQ 公布了一份它能检测的模型清单,并说会继续扩展到其他模型。FAQ 还说检测能力会随时间变化。以下是文档实际说了什么。

HumanPen 团队

· 12 分钟

简短回答

Turnitin 的 FAQ 公布了一份它能检测的 AI 写作模型清单。清单包括 GPT 系列、Gemini、Claude、LLaMA、Mistral、Deepseek、Nova、Grok 和 o1-mini 等。官方用的动词是"can detect content from." 清单以"and tools based on these LLMs as well." 结束。FAQ 还声明:"We will continue to expand our detection capabilities to other models in the future." FAQ 承认检测能力不是静态的:"As we iterate and develop our model further to better detect newer LLMs, it is likely that our detection capabilities will also change, affecting the AI percentage." GPT-5 这类具体模型是否在清单上取决于 FAQ 最后更新时间。FAQ 不承诺实时覆盖每个新模型发布,但承诺持续扩展。

模型清单说了什么

FAQ 提供了一份它能检测的模型清单。引导句用的是动词"can detect content from,"后面跟一个逗号分隔的模型列表。清单包括 GPT、Gemini、Claude 变体、LLaMA、Mistral、Deepseek、Nova、Grok 和 o1-mini。清单结尾是:

"and tools based on these LLMs as well."

下一句:"We will continue to expand our detection capabilities to other models in the future."

这意味着清单不是穷尽的。基于所列模型构建的工具也被声称可检测。扩展的承诺意味着新模型可能随检测模型的更新而被加入。Turnitin 能查出 Claude、Copilot、Gemini 吗把清单上点名的那几条过了一遍。

检测能力随时间变化

FAQ 明确承认检测不是固定的目标:

"As we iterate and develop our model further to better detect newer LLMs, it is likely that our detection capabilities will also change, affecting the AI percentage."

下一句:"However, for a submitted document, the AI percentage will change only if it's re-submitted again to be processed."

这有两个含义。第一,检测模型在更新以针对更新的 LLM。第二,已经打分的文档在模型变化时不会自动重新评估。如果你的论文在模型更新前被打分,分数反映的是当时的检测能力,不是当前的——所以Turnitin 复检后仍有标记:怎样比较两份报告,而不是只盯总分主张比报告而不是比分数。

检测器实际怎么工作

检测模型通过把文本切成重叠片段并分配概率分来处理:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."

模型不是基于 burstiness 或 perplexity 这类表面指标构建的。FAQ 声明:"Our model is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions." 下一句:"Instead, it learns statistical patterns from our training data."

同一 FAQ 还说:"Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers." 所以虽然模型不把 burstiness 或 perplexity 作为具名指标来计算,它确实基于从训练数据中学到的词概率模式来分类文本。

检测器可能漏掉什么

FAQ 承认误报和漏报之间的权衡:

"In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document."

FAQ 进一步说明:"For example, if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."

这意味着检测器宁可漏检一些 AI 文本也要避免误报。一个产生更接近人类词概率模式文本的新模型至少在检测模型更新覆盖它之前会更难被检测到。FAQ"continue to expand our detection capabilities"的承诺是制衡,但新模型发布和检测覆盖之间存在固有滞后。同样的滞后问题也出现在改写工具上:Turnitin 能检测出降 AI 工具吗

这对你意味着什么

总结一下我们讲的内容:

  • Turnitin 公布了一份它能检测的 AI 写作模型清单,包括 GPT、Gemini、Claude 等。
  • 官方动词是"can detect content from,"清单也包括基于这些 LLM 的工具。
  • FAQ 承诺会继续扩展检测到其他模型。
  • 检测能力随模型更新以针对更新的 LLM 而变化。
  • 已打分的文档在模型变化时不会自动重新评估。
  • 检测器以漏检一些 AI 文本为代价换取低误报率。

如果收到 Turnitin AI 报告并想处理被标记的段落,可以导入报告处理。符合条件时可以免费继续降 AI。

继续阅读