留学生降 AI 率:用二语英语写作需要知道什么
留学生用英语写学术论文,同时担心 AI 检测。Turnitin 的 FAQ 说训练数据包含了二语学习者以减少偏见。但检测器仍然按词概率模式分类,某些写作特征容易误判。本文解释这对你意味着什么。
HumanPen 团队
· 9 分钟
简短回答
Turnitin 的 FAQ 说模型的训练数据"took into account statistically under-represented groups like second-language learners, English users from non-English speaking countries, students at colleges and universities with diverse enrollments, and less common subject areas such as anthropology, geology, sociology, and others to minimize bias when training our model." 这是 Turnitin 自己的说法,不是独立验证过的结论。训练集纳入二语学习者意味着模型在训练时接触过二语写作。但检测器仍然通过对散文文本中的词概率模式做分类来工作。如果你的写作恰好和 AI 生成英语的统计模式有重叠,无论英语是不是你的母语,检测器都可能标记它。理解训练声明和检测机制之间的差距是应对标记的关键。
Turnitin 关于偏见的说法
FAQ 直接提到了偏见问题:
"While creating our sample dataset, we also took into account statistically under-represented groups like second-language learners, English users from non-English speaking countries, students at colleges and universities with diverse enrollments, and less common subject areas such as anthropology, geology, sociology, and others to minimize bias when training our model."
这是公司关于自己训练方法的声明。它说训练数据被设计为包含多样化的写作样本。但它没有给出任何数字、测试结果或外部对偏见水平的验证。声明说的是二语写作在训练中被代表了。声明没有说的是二语写作对误判免疫。模型在多样化数据上训练,但它仍然按词概率模式分类。为什么非英语母语的写作者更容易被 AI 检测器误判把这道缺口上的研究过了一遍。
检测器怎么读你的文本
检测流水线不管文本是谁写的都一样工作:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."
检测器不知道你的语言背景。它不根据你是不是英语母语者来调整分类。它读你词序列的统计特征并做分类。如果你的英语写作模式,可能因为正式教学强调某些学术短语和结构,恰好和模型与 AI 生成文本关联的模式有重叠,分数可能比你预期的更高。
误判和二语写作
FAQ 列出了容易误判的文本特征:
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."
下一句:"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
其中一些特征在二语学术写作中可能更常见。非母语者通常通过结构化模板和重复的学术短语学习英语,这可能产出结构变化较少的写作。转折词如"moreover"、"furthermore"、"in addition"被当作标准连接词教授,它们的重复使用可能产出 FAQ 描述的那种均匀结构。这不意味着二语写作总是被标记。它意味着和误判相关的特征可能和 instructed 二语写作中常见的模式有重叠。要不要因此改写法是另一个问题:为了不被标红,要不要改自己的写法。
你看不到自己的分数
还有一点留学生应该知道。FAQ 声明:"only instructors and administrators are able to see the indicator." 学生不能查看自己的 AI 分。FAQ 接着说:"However, with the PDF download feature, instructors can download and share the AI report with students." 如果你担心自己的分数,看到的唯一方法是请教师把 PDF 报告分享给你。你无法通过学生账号自己跑同样的检测。拿到 PDF 之后,下一步是怎么读一份 Turnitin AI 检测报告。
被标记后怎么办
总结一下我们讲的内容:
- Turnitin 称训练数据包含二语学习者以减少偏见,但未提供数字验证此声明。
- 检测器不管谁写的文本都读词概率模式,不管母语是什么。
- 误判易发特征如结构变化少,可能和二语学术写作常见的模式有重叠。
- 学生看不到自己的 AI 分。只有教师能看到,但他们可以分享 PDF。
- AI 分和相似度分是独立的测量。
如果你有 Turnitin 报告显示哪些段落被标记了,可以导入报告专门处理那些段落。符合条件时可以免费继续降 AI。
继续阅读