非母语写作被标记为 AI?文档怎么说
用二语写作已经够有挑战性了,还要担心 AI 标记。Turnitin 的 FAQ 说训练数据包含了二语学习者以减少偏见,但没有提供数字。检测器仍按词概率模式分类。非母语者常见的某些学术写作模式可能和误判易发特征有重叠。
HumanPen 团队
· 10 分钟
简短回答
非英语母语者可能在完全人写的文本上收到 AI 标记。Turnitin 的 FAQ 说模型的训练数据"took into account statistically under-represented groups like second-language learners, English users from non-English speaking countries"以减少偏见。这是公司自己的说法,未经独立验证。检测器仍然通过对散文文本中的词概率模式做分类来工作。教给非母语者的正式学术写作,强调结构化模板和标准转折词,可能产出结构变化少的文本。这个特征在 Turnitin 自己的误判易发文本类型清单上。被标记不意味着你的英语太好或太机械。它意味着你的文本统计特征和模型与 AI 关联的模式有重叠。
Turnitin 对训练数据多样性的说法
FAQ 直接提到偏见问题:
"While creating our sample dataset, we also took into account statistically under-represented groups like second-language learners, English users from non-English speaking countries, students at colleges and universities with diverse enrollments, and less common subject areas such as anthropology, geology, sociology, and others to minimize bias when training our model."
这是 Turnitin 对训练方法的描述。声明说的是二语写作在训练数据中被代表了。声明不包含的任何数字、测试结果或对模型在二语写作上表现的外部验证。FAQ 不公布按母语 vs 非母语分解的误报率。所以声明是用了多样化数据,但这个多样性努力的有效性无法从公开文档中独立测量。
检测器怎么读二语写作
检测流水线对所有文本都一样:
"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."
检测器不知道你的语言背景。它读词序列并做分类。如果你的写作模式,由正式英语教学塑造,产出的词概率序列和模型与 AI 生成文本关联的模式有重叠,分数可能比你预期的更高。具体是哪些模式,见困惑度与突发性之外,AI 检测器还在测什么。
与误判特征的重叠
FAQ 列出了误判易发文本:
"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."
下一句:"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
非英语母语教学通常强调结构化学术模板。学生学标准句子模式、正式转折词和一致的段落结构。这产出了 FAQ 描述的那种统一性的写作。重叠不是故意的。它是学术英语教学方式和检测器文本特征化方式的结果。要不要因此改写法是另一个问题:为了不被标红,要不要改自己的写法。
被标记了怎么办
FAQ 声明:
"Hence, we must emphasize that the percentage on the AI writing indicator should not be used as the sole basis for action or a definitive grading measure by instructors."
如果你的人写文本被标记了,你有理由和导师谈。你可以引用 FAQ 的误判易发文本描述,解释你的写作过程,并指出分数是一个数据点。你也可以引用训练数据多样性声明,同时诚实说明它不保证免于误判。被标记了,但确实是你自己写的:怎么准备申辩讲的是这套说法怎么搭起来。
这对你意味着什么
总结一下我们讲的内容:
- Turnitin 称训练数据包含二语学习者,但未提供数字验证有效性。
- 检测器不管语言背景都读词概率模式。
- 非母语者常见的正式学术写作可能和误判易发特征有重叠。
- FAQ 说分数不应作为采取行动的唯一依据。
- 被标记后,写作过程和 FAQ 自己的误判描述给了你讨论的依据。
如果你有 Turnitin 报告显示哪些段落被标记了,可以导入报告专门处理那些段落。符合条件时可以免费继续降 AI。
继续阅读