大模型生成参考文献常出错,新工具可显著提升准确性。
BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation
- 用真实数据构建基准,评估大模型生成引用的错误模式。
- 完整正确的引用仅50.9%,近期论文准确率比热门论文低27.7个百分点。
- 新工具clibib结合搜索与修正,提升准确率至91.5%且错误率极低。
大型语言模型结合网络搜索在科学出版自动化中日益普及,但其生成的BibTeX条目普遍存在字段级错误,包括遗漏、部分损坏、替换和虚构。我们构建了一个涵盖四个领域、三种引文层级(热门、低引、近期超截止)的931篇论文基准集,并提供版本感知的真值。三款前沿搜索模型(GPT-5、Claude Sonnet-4.6、Gemini-3 Flash)共生成约23,000个字段级观测结果。整体准确率为83.6%,但仅有50.9%的条目完全正确;从热门论文到近期论文,准确率下降27.7个百分点,揭示即使启用搜索,仍严重依赖参数化记忆。共现分析识别出两种失效模式:整条替换与孤立字段错误。我们提出开源工具clibib,实现确定性BibTeX检索,作为缓解机制。两阶段集成使准确率提升至91.5%(+8.0个百分点),完全正确条目达78.3%,回归率仅0.8%。分离搜索与修订的架构优于单阶段工具循环(0.8% vs. 4.8%),表明集成架构本身即具关键影响。我们已将clibib及配套代理技能以MIT许可证发布,旨在提升日益自动化的科研工作流中的引用准确性。
原文摘要 · Abstract (English)
Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive field-level errors stemming from omission, partial corruption, substitution, and hallucination. We construct a benchmark of 931 papers across four domains and three citation tiers---popular, low-citation, and recent post-cutoff---with version-aware ground truth. Three search-enabled frontier models (GPT-5, Claude Sonnet-4.6, Gemini-3 Flash) generate approximately 23,000 field-level observations. Overall accuracy is 83.6%, but only 50.9% of entries are fully correct; accuracy drops 27.7 pp from popular to recent papers, revealing heavy reliance on parametric memory even when search is available. Co-occurrence analysis identifies two failure modes: wholesale entry substitution and isolated field error. We present clibib, an open-source tool for deterministic BibTeX retrieval, as a mitigation mechanism. Two-stage integration raises accuracy to 91.5% (+8.0 pp) and fully correct entries to 78.3%, with a 0.8% regression rate. Separating search from revision yields larger gains and lower regression than single-stage tool loops (0.8% vs. 4.8%), demonstrating that integration architecture matters independently of model/tool capability. We release clibib and an accompanying agent skill under the MIT License to improve citation accuracy in increasingly automated scientific workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。