海报展示与 AI 检测:为什么分数不可靠

海报展示以项目符号、短标签和表格条目为主。Turnitin 的 FAQ 说模型只分析符合条件的散文句子,不对非散文内容做可靠检测。一篇帮助文章把表格也列入了不可靠范围。结果是百分比和高亮之间出现巨大差距。以下是检测机制在海报格式文档上的表现,以及为什么分数可能有误导性。

HumanPen 团队

· 22 分钟

简短回答

海报的写法和论文、期刊文章完全不同。它依赖项目符号、短词组和视觉元素。Turnitin 的文档说模型只分析符合条件的散文句子,不对非散文格式做可靠检测。当海报作为文档提交时,百分比可能只反映一小部分散文,而高亮几乎什么都没有。两者之间的差距是已知行为,不是故障。

项目符号不算符合条件的文本

FAQ 明确了什么才算合格文本

"This qualifying text includes only prose sentences, meaning that we only analyze blocks of text that are written in standard grammatical sentences and do not include other types of writing such as lists, bullet points (short non-sentence structures), or other non-sentence structures."

FAQ 还声明:

"The model does not reliably detect AI-generated text in the form of non-prose, or code, nor does it detect short-form/unconventional writing such as bullet points (short non-sentence structures)."

下一句解释了后果:

"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."

对海报来说,这个差距是极端的。海报上大部分文字是项目符号或短标签。如果你把海报作为文档提交并拿到高 AI 分,那个分数可能只基于少数几段符合条件的散文,而占海报大部分内容的项目符号对检测器是隐形的。

一篇帮助文章把表格也列入其中

一篇 Turnitin 帮助文章说模型的局限比 FAQ 列出的更广:

"The model does not reliably detect AI-generated text in the form of non-prose, such as poetry, scripts, or code, nor does it detect short-form/unconventional writing such as bullet points, tables, or annotated bibliographies."

该文章的下一句:

"This means that a document containing several different writing types would result in a disparity between the percentage and the highlights."

海报通常用表格展示结果、对比和时间线。如果表格被列为不可靠检测的范围,那么海报很大一部分内容就落在模型的可靠覆盖范围之外。你看到的百分比可能只是从表格和项目符号列表之间夹着的几段散文算出来的。

2023 年表格处理的更新

2023 年 8 月有一个重要更新。一份 release note 说:

"We are now able to process long-form prose text in tables."

下一句是:

"Resubmit to reprocess existing submissions that contain tables."

这意味着表格中包含长文散文句子(表格单元格里完整的语法句子)现在会被处理。表格中的短标签、数字和数据条目仍然是非散文,不在模型分析范围内。对海报来说这是部分改善。结果表格里如果每个格子都是完整句子,会被分析。如果只有数字和短标签,则不会。

检测机制与海报内容

FAQ 描述了提交文档的处理方式:

"When a paper is submitted to Turnitin, sentences from the submission are extracted and segmented into overlapping sections for prediction analysis. Each segment is classified by the AI detection model and given a value between 0 and 1, denoting the probability of the text being likely human or AI-generated. Each qualifying sentence within these segments inherits the segment's score."

只有符合条件的散文句子被提取。聚合后产生文档级分数:

"These sentence scores are further aggregated and used to compute the overall document AI writing score."

对海报来说,提取步骤可能只拉出很少的句子。海报上仅有的散文可能是摘要、简短引言和一句结论。其余全是项目符号和表格条目。分数从那少量散文中算出,意味着它反映的是很窄的样本。怎么读一份 Turnitin AI 检测报告也是从数字和页面对不上这一点开始讲的。

短文档的全有或全无问题

海报作为文档提交时篇幅很短。FAQ 对此有警告:

"In shorter documents where there are only a few hundred words, the prediction will be mostly 'all or nothing' because we're predicting on a single segment without the opportunity to overlap."

下一句:

"This means that some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated."

符合条件的散文只有三四百字的海报正好落在这个区间。片段太少,没有平均效果。一个高分片段就能把整篇文档标记为 AI 生成。FAQ 关于百分比和高亮之间差距的下一句在这里非常明显。你可能看到高百分比,但海报上几乎没有高亮文本,这属于AI 写作分不见了、对不上、下载不了里的一种情况。

误报风险与海报语言的重复性

FAQ 描述了容易触发误报的文本特征:

"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas."

海报散文倾向于重复。同样的关键发现出现在摘要、结果和结论里。因为空间有限,语言被压缩得统一,而这种统一恰恰接近困惑度与突发性之外,AI 检测器还在测什么里说的那个东西。FAQ 接着说:

"If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."

对海报来说这条建议格外重要。散文短、重复、结构统一。三个因素叠加,误报风险倍增。

这对你意味着什么

如果你的海报展示收到了 Turnitin AI 分,请记住以下几点:

  • 项目符号和短标签不是符合条件的文本。分数不反映它们。
  • 一篇 Turnitin 帮助文章说表格也不在可靠检测范围,不过 2023 年 8 月起表格中的长文散文会被处理。
  • 百分比和高亮之间的差距是混合格式文档的预期行为。
  • 条件散文少时会触发"全有或全无"预测,片段重叠不足。
  • 海报散文重复且统一,匹配 FAQ 描述的误报特征。
  • 高百分比但高亮少的情况应谨慎对待,不能作为定论。

如果想处理被标记的段落,可以导入 Turnitin 报告逐段修改。符合条件时可以免费继续降 AI。

继续阅读