arXiv:2503.23242cs.CLcs.AI2025-03中稿 · Computer magazine被引 1

首次实证检测到大模型生成文本在多语言虚假信息中的增长趋势。

Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation

  • 分析真实虚假信息数据集,追踪大模型文本出现变化
  • 发现ChatGPT发布后机器生成内容显著上升
  • 揭示跨语言、平台和时间的生成模式差异

大语言模型(LLMs)日益提升的文本生成能力及其多语言输出质量,引发了对其被用于制造虚假信息的担忧。尽管人类难以区分机器生成与人工撰写的文本,学界对这一威胁的影响仍存在分歧:部分观点认为风险被夸大,受限于自然生态系统的约束;另一些则指出特定‘长尾’场景面临被忽视的风险。本研究通过首个实证证据,揭示了最新真实世界虚假信息数据集中大模型文本的存在,记录了自ChatGPT发布以来机器生成内容的增长趋势,并展示了跨语言、平台及时间维度的关键模式。

原文摘要 · Abstract (English)

Increased sophistication of large language models (LLMs) and the consequent quality of generated multilingual text raises concerns about potential disinformation misuse. While humans struggle to distinguish LLM-generated content from human-written texts, the scholarly debate about their impact remains divided. Some argue that heightened fears are overblown due to natural ecosystem limitations, while others contend that specific "longtail" contexts face overlooked risks. Our study bridges this debate by providing the first empirical evidence of LLM presence in the latest real-world disinformation datasets, documenting the increase of machine-generated content following ChatGPT's release, and revealing crucial patterns across languages, platforms, and time periods.

虚假信息大模型检测多语言实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。