检测大模型伪造参考文献,保障学术引用可信度
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
- 设计多智能体流程,分步验证引用真实性
- 构建跨领域人工验证数据集,覆盖多种伪造类型
- 可规模化审计引用,适合科研人员与期刊编辑使用
科学研究依赖引用完整性,但大语言模型(LLMs)引入了关键风险:生成看似合理却无真实论文对应的虚构参考文献。手动验证已不可行,现有自动化工具也脆弱。我们提出CiteAudit,一个全面的基准与检测框架,用于识别幻觉引用。设计多智能体验证流程,将引用检查分解为元数据提取、记忆查询、网络检索和最终判断。为评估,构建大规模、人工验证的数据集,涵盖多个领域和多种伪造类型。实验表明,该框架在验证性能上优于现有最先进的LLMs和商业基线。本工作为大规模审计引用提供了必要基础设施,以维护学术话语的可信性。代码已开源。
原文摘要 · Abstract (English)
Scientific research relies on citation integrity, yet large language models (LLMs) have introduced a critical risk: fabricated references that appear plausible but correspond to no real publications. As manual verification becomes infeasible and existing automated tools remain fragile, we introduce CiteAudit, a comprehensive benchmark and detection framework for hallucinated citations. We design a multi-agent verification pipeline that decomposes citation checking into metadata extraction, memory lookup, web-based retrieval, and final judgment. To evaluate this, we construct a large-scale, human-validated dataset spanning diverse domains and hallucination types. Experiments demonstrate that our framework achieves superior verification performance over state-of-the-art LLMs and commercial baselines. Our work provides the necessary infrastructure to audit citations at scale and safeguard the trustworthiness of scholarly discourse. Code is available at https://github.com/shiiiikw/CiteAudit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。