系统综述与 AI 检测:方法学部分为什么被标记

系统综述遵循严格的方法学模板。方法学部分在不同研究间重复相同的动词和句式,匹配 Turnitin 描述的误报模式。检测器不是在测 burstiness 或 perplexity。它是词概率分类器。以下是文档关于分数如何产生以及如何解读的说法。

HumanPen 团队

· 21 分钟

方法学部分为什么匹配误报模式

系统综述要求标准化的方法学部分。你描述检索的数据库、使用的检索词、纳入和排除标准,以及筛选流程。语言必然重复,因为 PRISMA 清单要求特定的报告要素。Turnitin 的文档描述了产生误报的文本类型:

"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."

下一句:"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."

方法学部分命中全部三个标准,所以Turnitin 把我整章方法都标红了是很常见的抱怨。结构变化少,因为每篇系统综述报告相同的方法学步骤。措辞在各段落间重复("we searched"、"we identified"、"we screened")。文本大部分是对流程描述的复述,没有发展新观点。文档自己的指导意见说,当文本符合这些模式时要考虑百分比,而不是当作定论。

检测器如何计算分数

检测管道把合格文本分成片段并逐一评分:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score. Since segments overlap, some sentences may have multiple scores, which are then pooled into a single score. These sentence scores are further aggregated and used to compute the overall document AI writing score."

在系统综述中,方法学部分的重复句式可能在多个重叠片段上产生高分。这些分数汇总成一个高文档级百分比。真要动手改的话,方法与结果部分哪些可以改写,哪些必须逐字保持准确标出了不能碰的那几条线。一个 800 字的方法学部分可能拉高整篇 5000 字综述的分数,如果覆盖该部分的片段持续收到高概率值。

检测器实际测量什么

检测器不评估 burstiness、perplexity 或其他具名指标,这句话在Turnitin 到底用不用 perplexity 和 burstiness 来判定 AI里核过。文档明确说:

"Our model is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions."

下一句:"Instead, it learns statistical patterns from our training data."

同一来源解释了模型实际做什么:

"Our classifiers are trained to detect these differences in word probability and are adept at the particular word probability sequences of human writers."

检测器评估词概率序列。它检查一个片段中的词序列是否匹配从人类写作中学到的模式或从 AI 生成文本中学到的模式。在方法学部分,词概率序列由报告模板驱动。"we searched the following databases"和"studies were included if"这样的短语出现在成千上万篇论文中。这些序列可能更接近模型从训练数据中学到的统计模式,而不是某个人类作者的个性化选词。这种接近可以产生高概率分,而完全不涉及 AI。

什么算被分析的文本

检测器只处理合格文本

"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."

下一句:"This percentage is not necessarily the percentage of the entire submission."

系统综述通常包含表格(PRISMA 流程图、提取表、质量评估表)和结构化列表(纳入标准、排除标准)。表格和列表不是合格散文。百分比只覆盖散文句子。如果你的方法学部分主要是标准散文但结果部分主要是表格,百分比会偏向散文密集的部分。同一文档还指出:

"The model does not reliably detect AI-generated text in the form of non-prose, or code, nor does it detect short-form/unconventional writing such as bullet points (short non-sentence structures)."

下一句:"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."

系统综述是多格式文档。百分比与高亮之间的差距是预期行为。

短片段与全有或全无问题

系统综述的方法学部分有时被分成短小节。一个检索策略小节可能只有 200 字。文档对短文档发出警告:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

下一句:"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

对于短的方法学小节,检测器可能只在一个片段上预测。没有重叠来调节分数。一个完全由人写但遵循模板的小节可能得到全有或全无的高分。短篇幅消去了长文本所受益的平均效应。

Turnitin 已改进文档边界误报

2023 年,Turnitin 承认了文档边界的一个特定误报模式,并修改了检测逻辑:

"Since launch, we have observed a higher incidence of false positive detection in the first few or last few sentences of a document. Many times these sentences consist of introduction or conclusion content written in a generic way. As a result, we have changed our detection logic to help reduce these false positives."

下一句:"We also worked on making our segment boundaries detection more precise which could lead in some rare cases to change of boundaries compared with a previous version."

这是一项 2023 年的改进。文档的前几句和后几句(通常是引言和结论)产生了更高的误报率,因为它们往往以通用方式撰写。系统综述的引言和结论遵循模板(背景、空白、目标、发现摘要)。2023 年的逻辑变更针对了这一模式。片段边界精细化也改变了某些情况下文本被分成重叠片段的方式。

不要把分数当作唯一证据

Turnitin 的文档明确说明了分数的局限:

"Our AI writing detection model may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student."

下一句:"It takes further scrutiny and human judgment in conjunction with an organization's application of its specific academic policies to determine whether academic misconduct has occurred."

一篇系统综述的方法学部分得到高 AI 分数不是学术不端的证据。方法学部分匹配误报模式。检测器是词概率分类器,不是 burstiness 或 perplexity 评估器。分数是一个需要人类判断和机构政策的数据点。改卷人手上依据的那份指引,见AI 比例偏高时,厂商是怎么教老师处理的

这对你意味着什么

总结一下我们讲的内容:

  • 系统综述方法学部分匹配误报模式:结构变化少、措辞重复、复述式文本无新观点。
  • Turnitin 建议:"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."
  • 检测器是词概率分类器,不是 burstiness 或 perplexity 评估器。它从训练数据中学习统计模式。
  • 只有合格散文被分析。表格、列表和要点被排除,造成百分比与高亮之间的差距。
  • 短方法学小节因单片段预测面临全有或全无的评分问题。
  • 2023 年,Turnitin 改进了检测逻辑以减少文档边界(引言和结论)的误报。
  • 分数不应作为不利处分的唯一依据。它需要人类判断。

如果收到 Turnitin AI 报告并想处理被标记的段落,可以导入报告处理。符合条件时可以免费继续降 AI。

继续阅读