Turnitin AI 检测中的"合格文本"是什么?详解
当 Turnitin 报告一个 AI 写作百分比时,那个数字不覆盖你的整篇提交。模型只分析"合格文本",也就是标准语法形式的散文句子。列表、项目符号和非句子结构被排除在外。这解释了为什么百分比和高亮段落有时不匹配。以下是文档关于什么算合格、什么不算、以及为什么重要的实际说明。
HumanPen 团队
· 15 分钟
"合格文本"是什么意思
Turnitin 的 AI 检测模型不会分析文档中的每一个字。它作用于一组特定的文本,称为"合格文本"。文档对此有明确定义。
"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."
所以模型寻找的是散文:用标准语法句子组织的文本块。任何不满足这个标准的文本都被排除在分析之外。项目符号、编号列表、不是完整句子的表格标题和其他非散文结构都不属于 AI 检测的范围。
模型自己的文档承认了这个限制:"The model does not reliably detect AI-generated text in the form of non-prose, or code, nor does it detect short-form/unconventional writing such as bullet points (short non-sentence structures)."
如果一个文档列表很多、散文很少,AI 检测模型可处理的材料就少,报告的分数只反映散文部分。下面还有一条底线:文件要求至少 300 词散文才会生成报告,见一整本学位论文,Turnitin 查得了吗。
百分比和高亮之间的差异
Turnitin AI 检测报告最令人困惑的一点是,百分比和高亮段落看起来在讲不同的故事。文档解释了原因。
"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."
"This percentage is not necessarily the percentage of the entire submission."
如果你的文档一半是散文、一半是项目符号,AI 检测模型标记了其中 40% 的散文为 AI 生成,报告可能显示 40% 的 AI 分数。但这个 40% 指的是合格散文,不是整个文档。不了解这一点的人可能认为整篇提交有 40% 是 AI 生成的,而实际占整个文档的比例更低。
这个差异不是 bug。它是模型范围的产物。理解它有助于准确阅读报告,页面上其余部分怎么看,见怎么读一份 Turnitin AI 检测报告。
模型如何处理文本
要理解合格文本为什么重要,需要知道模型用它做了什么。文档描述了机制:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."
模型提取句子,将它们分段为重叠的区间,然后对每个区间分类。每个区间得到一个 0 到 1 之间的概率分数。只有合格散文句子会经过这个流程。非散文内容不会被分段或评分,因为模型不把它当作可分析的文本。
这就是为什么"合格文本"的范围直接决定了百分比代表什么。分数是区间级概率分数的汇总,只有合格散文句子参与了汇总。
表格算,参考文献不算
两种特定内容类型有自己的规则。
表格:"We are now able to process long-form prose text in tables." 所以表格单元格内的散文,如果用标准语法句子写成,现在属于合格文本。Resubmit to reprocess existing submissions that contain tables. 如果你有一篇表格中包含有意义散文的文档,且在那次更新之前已提交,AI 分数可能不反映表格内容,除非重新提交。
参考文献:"We have fixed a bug that was occasionally highlighting AI writing within references listed in a bibliography. Bibliographies are now excluded when processing the AI writing report." Resubmit to reprocess existing submissions that contain highlighted reference sections. 如果你之前的报告在参考文献列表中显示了被标记的文本,那是一个 bug,重新提交会把这些部分排除在分析之外。相似度那一侧的排除项管到哪一步,见Turnitin 为什么把参考文献和引用标出来。
这两条规则意味着合格文本的定义不是静态的。它已经被调整,纳入了表格散文,排除了参考文献。这也是文本没改、重新提交后 Turnitin AI 分数会变的原因之一。
合格散文中的误报
即使在合格散文内部,某些类型的写作也更容易出现误报。文档指出了具体的模式。
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."
"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
所以如果你的散文结构单调、重复相同措辞,或者只是改写而没有发展新观点,模型可能把它标记为 AI 生成,即使它是人类写的。这不是合格文本过滤器的问题,而是AI 检测器还在测什么里那套东西按训练目标正常工作的结果。这是在确实参与分析的散文内部的一个独立问题。
这对你意味着什么
如果你想理解为什么你的 Turnitin AI 报告呈现出现在的样子,以下是总结:
- 只有散文句子算合格。 列表、项目符号和非句子结构被排除。模型聚焦于标准语法句子。
- 百分比只覆盖合格散文。 它不是整个提交的百分比。这造成数字和高亮之间的差异。
- 表格散文现在会被处理。 表格中的长篇散文算作合格文本。Resubmit to reprocess existing submissions that contain tables.
- 参考文献被排除。 一个曾经标记参考文献的 bug 已被修复。Resubmit to reprocess existing submissions that contain highlighted reference sections.
- 注意误报。 结构扁平、重复或改写的散文更容易被标记。阅读百分比时需要考虑这一点。
符合条件时可以免费继续降 AI。
继续阅读