arXiv:2605.07723cs.DLcs.AI2026-05被引 8

发现大模型在论文中生成大量虚构引用,2025年至少14万条,影响科学可信度与公平性。

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

  • 通过验证数百万论文的引用,发现大模型导致虚假参考文献激增。
  • 2025年仅arXiv等平台就存在至少146,932条虚构引用,且集中在快速采用AI的领域。
  • 错误引用多出现在小团队和青年学者论文中,且倾向奖励已有名望的男性学者。

大型语言模型(LLMs)在多种场景下生成看似合理却虚假的信息,但其真实世界的影响程度仍不清楚。本研究利用可验证的科学引文作为审计对象,分析了来自arXiv、bioRxiv、SSRN和PubMed Central的250万篇论文中的1.11亿条参考文献。结果显示,在大模型广泛使用后,不存在的参考文献数量显著上升,保守估计2025年即有146,932条幻觉引用。这些错误分布广泛,但在人工智能快速渗透的领域尤为突出;具有AI辅助写作语言特征的论文、小型及早期职业作者团队中问题更严重。同时,幻觉引用倾向于将学术贡献归于已有声望的男性学者,可能加剧科研评价中的不平等。预印本审核与期刊发表流程仅捕获部分错误,表明幻觉内容的传播已超过现有防控能力。综合来看,大模型幻觉正大规模渗入知识生产,威胁未来科学研究的可靠性与公平性。

原文摘要 · Abstract (English)

Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and consequences of this hallucination problem remain poorly understood. Here we leverage a uniquely verifiable object - scientific citations - to audit 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central. We find a sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. These errors are diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams. At the same time, hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition. Preprint moderation and journal publication processes capture only a fraction of these errors, suggesting that the spread of hallucinated content has outpaced existing safeguards. Together, these findings demonstrate that LLM hallucinations are infiltrating knowledge production at scale, threatening both the reliability and equity of future scientific discovery as human and AI systems draw on the existing literature.

大模型幻觉虚假引用科研诚信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。