Turnitin 如何处理数学公式和非散文内容
数学和工程专业的同学经常想知道公式和方程是否影响 AI 分。Turnitin 的 FAQ 说模型只分析符合条件的散文句子,不可靠检测非散文内容。FAQ 还说 Turnitin 不追求代码检测。以下是什么算符合条件的文本、什么不算。
HumanPen 团队
· 13 分钟
简短回答
数学公式、方程和代码块不被 Turnitin 的 AI 检测器分析。FAQ 声明:"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures." 模型还"does not reliably detect AI-generated text in the form of non-prose, or code." FAQ 补充:"In addition, we are not pursuing ChatGPT code detection at this time." 这意味着你的公式和方程不直接贡献 AI 分。分数只反映文档中的散文部分。如果你的论文主要是公式加少量解释文字,AI 检测器评估的文本量很小,这会降低分数的意义。
什么算符合条件的文本
FAQ 定义了模型检查的范围:
"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."
下一句:"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."
对于数学或工程论文,这意味着公式本身不被评估。只有它们周围的散文句子、解释、引言和讨论才接受 AI 检测。如果你有一整页公式加两句解释,检测器评估的就是那两句话。轮到改写时,问题就反过来了:改写工具会把引用、表格和公式改成什么样。
代码和非散文内容
FAQ 对非散文内容说得很直接:
"The model does not reliably detect AI-generated text in the form of non-prose, or code, nor does it detect short-form/unconventional writing such as bullet points (short non-sentence structures)."
FAQ 还声明:"In addition, we are not pursuing ChatGPT code detection at this time."
这对 STEM 论文有实际含义。如果你用 AI 工具生成代码或公式,AI 检测器不是为抓那个设计的。检测器评估的是周围的散文,不是代码或数学本身。
机制如何在散文上工作
检测模型通过把符合条件的文本切成重叠片段来处理:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated."
片段重叠,意味着句子可以接收多个分数并汇总在一起。在公式密集、散文稀疏的论文中,重叠优势减弱,因为可供切分的散文句子更少。这会降低分数的稳定性,因为模型可用的数据点更少。
表格和数据怎么办?
2023 年 8 月的一条 release note 描述了模型如何处理表格:
"We are now able to process long-form prose text in tables."
下一句:"Resubmit to reprocess existing submissions that contain tables."
这意味着如果你的表格包含长文散文句子,那些句子现在会被分析。但表格中的原始数据、数字和短标签不是散文句子,不在符合条件的文本定义内。参考书目也被排除:同一 release note 说"Bibliographies are now excluded when processing the AI writing report." 下一句同样要求重新提交已有的提交才能生效——这也是为什么在旧报告上看起来仍然是Turnitin 把参考文献标红了。
AI 分和相似度分是独立的
一个常见的困惑是数学内容在相似度检查中被标记是否也影响 AI 分。不会。FAQ 声明:
"The Similarity score and the AI writing detection percentage are completely independent and do not influence each other."
如果你的公式在相似度数据库中匹配到某个来源,那影响的是相似度分,不是 AI 分。AI 检测器只评估散文是否有 AI 生成的模式,不管相似度检查发现了什么;这里要避开的错误是两种分数不要混着看。
这对你意味着什么
总结一下我们讲的内容:
- 数学公式和方程不被 AI 检测器分析。只有符合条件的散文句子才被分析。
- FAQ 说模型不可靠检测非散文或代码,且不追求代码检测。
- 包含长文散文的表格现在会被处理,但原始数据和短标签不会。
- 参考书目在 AI 检测处理中被排除。
- AI 分和相似度分完全独立,互不影响。
- 公式密集、散文稀疏的论文可供模型评估的文本更少,可能导致分数不够稳定。
如果收到 STEM 论文的 Turnitin 报告并想处理被标记的散文段落,可以导入报告处理。符合条件时可以免费继续降 AI。
继续阅读