arXiv:2410.08044cs.CL2024-10被引 37

检测发现新维基文章中5%为AI生成,质量低且易偏颇。

The Rise of AI-Generated Content in Wikipedia

  • 用双工具检测近期维基页面,对比GPT-3.5前后差异。
  • 英语维基新文章超5%被标记为AI生成,德法意语种较低。
  • 被标记文章多为低质、自夸或立场偏颇内容。

人工智能生成内容在主流信息源中的兴起引发了问责、准确性及偏见放大的严重担忧。除了直接影响读者外,这类内容的广泛存在也对基于大规模互联网数据训练语言模型的长期可行性提出质疑。本文使用GPTZero(专有检测工具)和Binoculars(开源替代方案),对近期创建的维基百科页面进行检测,以建立人工智能生成内容的下限估计。两个检测器均显示,相较于GPT-3.5发布前的文章,近期文章中人工智能生成内容显著增加。在将误报率控制在预GPT-3.5文章上的1%阈值下,超过5%的新英语维基百科文章被标记为人工智能生成,德语、法语和意大利语文章的比例则更低。被标记的文章通常质量较低,常带有自我推广性质,或在争议性话题上表现出特定倾向。

原文摘要 · Abstract (English)

The rise of AI-generated content in popular information sources raises significant concerns about accountability, accuracy, and bias amplification. Beyond directly impacting consumers, the widespread presence of this content poses questions for the long-term viability of training language models on vast internet sweeps. We use GPTZero, a proprietary AI detector, and Binoculars, an open-source alternative, to establish lower bounds on the presence of AI-generated content in recently created Wikipedia pages. Both detectors reveal a marked increase in AI-generated content in recent pages compared to those from before the release of GPT-3.5. With thresholds calibrated to achieve a 1% false positive rate on pre-GPT-3.5 articles, detectors flag over 5% of newly created English Wikipedia articles as AI-generated, with lower percentages for German, French, and Italian articles. Flagged Wikipedia articles are typically of lower quality and are often self-promotional or partial towards a specific viewpoint on controversial topics.

AI内容检测维基百科语言模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。