arXiv:2602.06718cs.CRcs.AI2026-02被引 16

发现大模型写论文时会大量虚构参考文献,威胁学术可信度。

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

  • 构建开源框架验证引用真实性,系统检测大模型生成的虚假引用。
  • 13个大模型引用错误率14.23%至94.93%,2025年无效引用增长80.9%。
  • 超七成研究者用AI写作,八成审稿人不查参考文献,需集体应对。

引用是信任科学结论的基础;若引用无效或伪造,这种信任将崩塌。随着大语言模型(LLMs)在学术写作中日益普及,这一风险加剧:大模型存在虚构引用(“幽灵引用”)的倾向,对引用有效性构成系统性威胁。为量化该风险,我们开发了 \\_citeb,一个开源的大规模引用验证框架,并通过三项互补实验开展全面研究。首先,在多个研究领域对13个大模型进行引用生成任务基准测试,发现所有模型的引用幻觉率介于14.23%至94.93%之间。其次,分析2020–2025年间来自人工智能/机器学习与安全领域会议的56,381篇论文中的220万条引用,发现1.07%的论文包含无效引用,且2025年该比例较往年上升80.9%。第三,对97名研究人员进行调查,结果显示87.2%的研究者在其工作流程中使用基于AI的工具,76.7%的审稿人未彻底检查参考文献,74.5%的人认为同行评审难以识别引用错误。基于这些发现,我们认为幽灵引用已成为威胁学术诚信的系统性问题,呼吁学界协同应对。

原文摘要 · Abstract (English)

Citations provide the basis for trusting scientific claims; when they are invalid or fabricated, this trust collapses. With the advent of Large Language Models (LLMs), this risk has intensified: LLMs are increasingly used for academic writing, but their tendency to fabricate citations (``ghost citations'') poses a systemic threat to citation validity. To quantify this threat, we develop \citeb, an open-source framework for large-scale citation verification, and conduct a comprehensive study of citation validity in the LLM era through three complementary experiments. First, we benchmark 13 LLMs on citation generation task in various research domains, finding that all models hallucinate citations at rate from 14.23\% to 94.93\%. Second, we analyze 2.2 million citations from 56,381 papers at AI/ML and Security venues (2020--2025), finding that 1.07\% of papers contain invalid citations, with an 80.9\% increase in 2025. Third, we survey 97 researchers, finding that 87.2\% use AI-powered tools in their workflows, 76.7\% of reviewers do not thoroughly check references, and 74.5\% view peer review as ineffective at catching citation errors. Based on these findings, we argue that ghost citations represent a systemic threat to academic integrity, and call for coordinated efforts from community to address this challenge.

大模型引用验证学术诚信幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。