约束提示导致大模型编造看似真实的引用,影响学术可信度。
Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes
- 在五种提示策略下测试四模型引用真实性
- 17,443条引用中真实存在率最高仅47.5%
- 编造引用占未验证结果的三成以上,需事后核验
大模型常生成看似合法的虚假参考文献,尤其在部署约束提示下。本研究基于144个陈述(24个来自软件工程与计算机科学领域),使用确定性验证流程(Crossref + Semantic Scholar)评估了两款专有模型(Claude Sonnet、GPT-4o)和两款开源模型(LLaMA~3.1-8B、Qwen~2.5-14B)在五种提示范式下的表现:基础、时间窗口、综述式广度、保密政策及组合策略。共生成17,443条引用,引用存在率最高为0.475;时间窗口与组合条件下降最显著,但输出格式仍合规。未决结果占比达36%-61%,100条审计显示大量未决项为虚构。研究呼吁在进入软件工程文献综述或工具链前进行事后引用验证。
原文摘要 · Abstract (English)
LLMs are increasingly used to draft academic text and to support software engineering (SE) evidence synthesis, but they often hallucinate bibliographic references that look legitimate. We study how deployment-motivated prompting constraints affect citation verifiability in a closed-book setting. Using 144 claims (24 in SE&CS) and a deterministic verification pipeline (Crossref + Semantic Scholar), we evaluate two proprietary models (Claude Sonnet, GPT-4o) and two open-weight models (LLaMA~3.1-8B, Qwen~2.5-14B) across five regimes: Baseline, Temporal (publication-year window), Survey-style breadth, Non-Disclosure policy, and their combination. Across 17,443 generated citations, no model exceeds a citation-level existence rate of 0.475; Temporal and Combo conditions produce the steepest drops while outputs remain format-compliant (well-formed bibliographic fields). Unresolved outcomes dominate (36-61%); a 100-citation audit indicates that a substantial fraction of Unresolved cases are fabricated. Results motivate post-hoc citation verification before LLM outputs enter SE literature reviews or tooling pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。