arXiv:2607.18360cs.CRcs.AI2026-07

提出首个针对大模型引文验证器的系统性故障诊断基准,揭示误报率是部署关键瓶颈。

HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers

论文配图:HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
图 1 · 摘自论文原文
  • 构建包含2526条引文的基准HALLMARK,覆盖14类幻觉、三难度层级和六项诊断测试
  • 发现误报率远高于召回率,决定验证器能否实际部署,尤其在真实投稿率下
  • 多数大模型对训练截止后论文误报严重,仅最新模型保持低误报率,适合学术场景

大型语言模型(LLMs)常用于撰写文献综述和辅助学术写作,导致伪造参考文献风险上升:GPTZero在NeurIPS 2025的录用论文中发现了53篇存在幻觉引文。规则与基于LLM的验证器相继出现,但缺乏统一基准进行比较和详细故障诊断。为此,我们提出HALLMARK(幻觉检测基准):包含2,526条BibTeX条目,涵盖14种幻觉类型、三个难度层级、每条6个诊断子测试,并采用抗污染的留出划分。我们在该基准上评估了基于DOI查询的基线、前沿零样本大模型、工具增强代理及我们自研的规则式共设计验证器bibtex-updater。结果一致显示:是否可部署取决于误报率(FPR),而非召回率。具体表现为:代理式查询虽提升召回率,但显著增加误报;在真实投稿率下,误报率数量级差异(而非召回率)决定了验证器标记是否为有效信号还是噪声;大多数大模型对训练截止后发表的论文过度标记,仅两个最新截止模型维持近分布水平的误报率(此现象作为描述性信号报告,因可能受召回影响)。因此,误报率是部署瓶颈,但未被发现的伪造引用对科学记录代价更高。

原文摘要 · Abstract (English)

Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set. Rule- and LLM-based verifiers are emerging, but no shared benchmark compares them and gives detailed failure diagnostics. We close that gap with HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split. On it we evaluate a DOI-lookup baseline, frontier LLMs zero-shot, tool-augmented agents, and our own rule-based, co-designed verifier bibtex-updater. Across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deployable. HALLMARK makes it concrete through three failure modes: agentic lookups buy recall but inflate false positives; at a venue-realistic base rate, the order-of-magnitude spread in false-positive rates (FPRs) -- not recall -- governs whether a verifier's flags are mostly true catches or mostly noise; and most LLMs over-flag papers published past their training cutoff, where only the two latest-cutoff models hold their false-positive rate near in-distribution levels (a signal we report as descriptive, since it is confounded with possible recall of those entries). Thus FPR is the deployment bottleneck, but an undetected fabrication remains the costlier error for the scientific record.

大模型验证引文幻觉误报率基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。