大学正在放弃 AI 检测吗
越来越多大学在 AI 检测分数面前加上了程序性保障措施。我们来看 Turnitin 自己的文档对"正确用法"说了什么,为什么误报率在万人规模下不是小数字,以及为什么这个工具短期内不会消失。
HumanPen 团队
· 23 分钟
简短回答
大学并没有在"放弃 AI 检测"这个意义上真正抛弃这项技术。它们放弃的是"把 AI 分数当作指控学术不端的唯一依据"这一政策。转变的方向是从"分数就是证据"走向"分数是需要进一步审查的一条线索"。这个转变不是在抵触 Turnitin,而是在跟上 Turnitin 自己文档里早就写明的建议。
Turnitin 的 AI 写作报告使用说明 里有一句话说得很直白。厂商原话是:分数 "should not be used as the sole basis for adverse actions against a student." 那些要求在指控学生之前收集额外证据的大学,遵从的正是厂商自己的指引。
Turnitin 的文档对"正确用法"到底说了什么
理解这股趋势最关键的一句话来自 Turnitin 的 AI 写作报告使用说明。Turnitin 说:
"Our AI writing detection model may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student."
紧跟着的下一句同样必须读完:
"It takes further scrutiny and human judgment in conjunction with an organization's application of its specific academic policies to determine whether academic misconduct has occurred."
这两句话合在一起,正好描述了各大学正在采用的策略。分数触发审查,但审查不等于定论。学生不会因为一个数字就被告发,而是因为一个数字进入审查流程,审查中要运用人工判断和机构自身政策,才能得出是否构成不端的结论。
Turnitin 另一页独立指引 How should I review the AI Writing report? 也在重复同一个框架:
"It is not meant to provide definitive answers in isolation. More important than any tool is the educator who sees the score and makes decisions balancing this information with their personal knowledge of their students, their work, and institutional policy."
该页的下一句把预期用途说得更清楚:
"When educators look at the AI writing score and utilize it as a single data point rather than a definitive response, then it is being used as intended."
"a single data point rather than a definitive response",就是这一句。采用这个定位的大学并没有在跟 Turnitin 对着干,它们做的是厂商说这套工具本该被怎么用的事。
误报率在大学规模下的数学账
Turnitin 说误报率低于 1%,但细节很重要。FAQ 原文是:
"We strive to maximize the effectiveness of our detector while keeping our false positive rate - incorrectly identifying fully human-written text as AI-generated - under 1% for documents with over 20% of AI writing."
下一句用更直白的方式重述了这个比例:
"In other words, we might flag a human-written document as AI-written for one out of every 100 fully-human written documents."
"for documents with over 20% of AI writing" 这个限定条件不能丢。低于 1% 的目标针对的是 AI 写作占比超过 20% 的文档。重述句把这个限定抹掉了,但原始声明没有。
即便按 100 篇里 1 篇算,这个账放到万人规模上就不好看了。一个每学期处理一万篇论文的大学,可能有大约 100 篇完全由人写成的文档被标记为 AI 生成。一个处理 7.5 万篇论文的大学,这个数字是 750。这不是抽象数字。每一个数字背后都是一个因为工具里的一个百分比而面临不端指控的学生,而厂商自己说了这个百分比不能当唯一依据用。
(范德堡大学停用 AI 检测器的决定在另一篇文章里已经讲过。这里关注的是更大的趋势,不是某一个学校。)
促使大学谨慎对待的两个已知精度缺口
检测器的已知限制不止两个,其中两条可以解释为什么大学在加程序而不是撤工具。
第一,短文档对模型来说很难办。FAQ 说:
"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."
下一句:
"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."
如果一门课的作业是几百字的短评或反思,这是个很实际的问题。一篇 300 字的作业哪怕只有一部分借助了 AI,也可能被整篇标成"100% AI 生成"。
第二,厂商承认在召回率和精确率之间做了取舍。FAQ 说:
"In order to maintain this low rate of 1% for false positives, there is a chance that we might miss some AI written text in a document. We're comfortable with that since we do not want to incorrectly highlight human-written text as AI-written. For example, if we identify that 50% of a document is likely written by an AI tool, it could contain as much as 65% AI writing."
也就是说检测器在设计上偏保守。它宁可漏掉一些 AI 文本,也不想冤枉人写的文字。担心误报的大学能从这种设计取向里获得一些安心。但反过来,这也意味着一份干净的分数并不证明文档里没有 AI,只说明工具没越过它的标记阈值。
工具本身并没有停下
有一个事实让"大学放弃 AI 检测"这个叙事变得更复杂:厂商自己没有停掉这个工具。FAQ 说:
"In July 2026, we updated our model architecture to consolidate a multi-model ensemble into a single model. This update improves and simplifies the AI writing report, maintaining a less than 1% false positive rate."
模型仍在持续开发。2026 年 7 月的这次合并简化了架构,同时维持了误报率目标。大学的政策调整发生在一个厂商仍在投入的工具上。技术没有消失,机构对它的应对方式在变成熟。
政策走向:从"唯一依据"到"辅证之一"
我们在各机构观察到的,不是"用检测"和"不用检测"的二元选择。模式更具体一些。大学的政策大致在三个阶段之间移动。
第一阶段:分数被当作不端认定成立的充分证据。这是"唯一依据"路线。Turnitin 自己的文档说工具不该这么用。
第二阶段:分数触发一个需要更多证据的审查流程。审查可能包括和学生谈话、对比以往作业、检查草稿历史。这是"辅证"路线。它和 FAQ 里说的分数 "should not be used as the sole basis" 以及 "it takes further scrutiny and human judgment" 一致。
第三阶段:机构停用 AI 检测器。这比较少见。范德堡的决定是被引用最多的例子,但大多数重新审视过政策的机构停在了第二阶段,没有走到第三阶段。
趋势指向第二阶段。大学保留工具,加上程序保障。它们没有放弃检测,它们在做的是要求把分数当成 Turnitin 自己说的那种东西来对待:一个数据点,不是一个定论。
对学生和教师的实际意义
如果你是学生,实际结论是:你论文上的高分 AI 标记不会自动等于指控。在越来越多的学校里,分数开启的是一次对话,不是一个判决。保留你的草稿,保留版本历史,被问到时能说清自己的写作过程就够了。
如果你是教师,结论是 Turnitin 自己的文档已经为这种政策转变提供了背书。FAQ 说分数 "should not be used as the sole basis for adverse actions." 审查指引页说它是 "a single data point rather than a definitive response." 你不需要把"辅证"政策包装成偏离厂商建议的做法,它就是厂商的建议。
如果你有一篇论文已经被标记,想在提交前修改,我们可以帮上忙。
继续阅读