研究论文引用如何影响大模型判断,发现带引用就更易胡说八道。
Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models

- 构建22万条提示的跨领域基准,分离真假陈述与真假引用。
- 有引用时幻觉率上升3-22个百分点,最高达77%。
- 伪造引用+真陈述最危险,法律类模型最抗干扰。
大语言模型在依赖引用的场景中应用日益广泛,但引用本身对模型行为的影响(脱离事实内容)尚不明确。本文提出AuthorityBench,一个包含220,564个提示的多领域基准,通过完全平衡的2×2因子设计,分离陈述真假与引用真假,在通用知识、科学、法律和医学四个领域展开研究。实验采用40种提示模板、四级期刊声望等级及带国家编码的作者姓名数据集,评估7个模型在12个结构化研究问题上的表现。结果表明,无论引用真实与否,存在引用均显著提升幻觉率,尤其在虚假引用搭配真实陈述时,幻觉率提升3至22个百分点,通用知识领域高达35%至77%;法律类陈述则相对稳健。期刊声望与作者身份特征影响微弱。所有数据与代码已开源。
原文摘要 · Abstract (English)
Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual content remains poorly understood. We introduce AuthorityBench, a 220,564-prompt multi-domain benchmark that isolates how citation-based authority signals influence epistemic behavior in LLMs. The benchmark uses a fully balanced 2x2 factorial design crossing claim veracity with citation veracity, the first to do so, across four domains (general knowledge, science, law, and medicine), with controlled variation over 40 prompt templates, four venue prestige tiers, and a country-coded author name dataset. Evaluating seven models on 12 structured research questions, we find that citation presence, whether real or fabricated, consistently increases hallucination rates relative to a no-citation baseline. The effect is strongest when fabricated citations accompany true claims, raising hallucination rates by 3 to 22 percentage points and reaching 35 to 77% in the general knowledge domain, while legal claims are comparatively robust and venue prestige and author demographics show negligible impact. All datasets and evaluation code are available at: https://github.com/floating-reeds/AuthorityBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。