arXiv:2603.03299cs.CL2026-03被引 1

检测大模型虚构参考文献,发现引用错误率高达56.8%。

How LLMs Cite and Why It Matters: A Cross-Model Audit of Reference Fabrication in AI-Assisted Academic Writing and Methods to Detect Phantom Citations

  • 测试10个主流大模型在4个领域生成6.9万条引用,验证其真实性
  • 引用幻觉率在11.4%到56.8%之间,受模型、领域和提示词影响大
  • 提出双滤法与轻量分类器,可高效识别虚构引用,适合学术写作审核

大型语言模型(LLMs)存在虚构学术引用的问题,但其在不同厂商、领域及提示条件下的表现尚不明确。本文开展了目前规模最大的引用幻觉审计,对10个商用部署的LLM在四个学术领域中生成的69,557条引用进行了跨三大学术数据库(CrossRef、OpenAlex、Semantic Scholar)的验证。结果显示,幻觉率跨度达五倍(11.4%至56.8%),且受模型、领域和提示方式显著影响。研究还发现,未被提示时所有模型均不会自发生成引用,表明幻觉为提示诱导而非内在特性。本文提出两种实用过滤策略:多模型共识(超过3个模型引用同一文献时准确率达95.6%,提升5.8倍)和同提示重复(重复超过2次时准确率达88.9%)。此外,研究揭示了模型迭代中改进并非必然,而同家族模型容量扩大可降低幻觉率。最后,基于书目字符串特征训练的轻量级分类器,在交叉验证中达到AUC 0.876,LOMO泛化测试中达0.834,无需外部数据库查询,可作为推理阶段的预筛选工具。

原文摘要 · Abstract (English)

Large language models (LLMs) have been noted to fabricate scholarly citations, yet the scope of this behavior across providers, domains, and prompting conditions remains poorly quantified. We present one of the largest citation hallucination audits to date, in which 10 commercially deployed LLMs were prompted across four academic domains, generating 69,557 citation instances verified against three scholarly databases (namely, CrossRef, OpenAlex, and Semantic Scholar). Our results show that the observed hallucination rates span a fivefold range (between 11.4% and 56.8%) and are strongly shaped by model, domain, and prompt framing. Our results also show that no model spontaneously generates citations when unprompted, which seems to establish hallucination as prompt-induced rather than intrinsic. We identify two practical filters: 1) multi-model consensus (with more than 3 LLMs citing the same work yields 95.6% accuracy, a 5.8-fold improvement), and 2) within-prompt repetition (with more than 2 replications yields 88.9% accuracy). In addition, we present findings on generational model tracking, which reveal that improvements are not guaranteed when deploying newer LLMs, and on capacity scaling, which appears to reduce hallucination within model families. Finally, a lightweight classifier trained solely on bibliographic string features is developed to classify hallucinated citations from verified citations, achieving AUC 0.876 in cross-validation and 0.834 in LOMO generalization (without querying any external database). This classifier offers a pre-screening tool deployable at inference time.

大模型引用幻觉学术写作可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。