arXiv:2602.01031cs.AIcs.CL2026-02被引 6

构建多轮对话幻觉评测基准,揭示大模型事实错误率仍超三成。

HalluHard: A Hard Multi-Turn Hallucination Benchmark

  • 设计950个高风险领域问题,要求答案带可验证引用。
  • 引入网页搜索迭代检索机制,自动验证引用是否支持内容。
  • 发现即使使用网络搜索,顶级模型幻觉率仍达30%,且受推理能力影响。

大型语言模型在多轮对话中常生成看似合理但缺乏依据的事实性错误,且随着上下文累积和早期错误传播而加剧。我们提出霍鲁哈德(HalluHard),一个包含950个种子问题的挑战性多轮幻觉评估基准,覆盖法律案件、科研问题、医疗指南和编码四大高风险领域。通过要求生成内容中的事实陈述必须附带内联引用,实现对事实依据的量化评估。为支持开放场景下的可靠判断,我们设计了一套迭代式证据检索流程:自动执行网络搜索,获取、过滤并解析全文资料(包括PDF),以验证引用是否真正支持生成内容。在一系列前沿闭源与开源模型上测试发现,即便启用网络搜索,幻觉现象依然严重——最强配置(Opus-4.5 + 网络搜索)的幻觉率约为30%;内容与事实脱节的问题持续普遍存在。最后,我们证明幻觉行为受模型容量、对话轮次位置、有效推理能力及知识类型共同影响。

原文摘要 · Abstract (English)

Large language models (LLMs) still produce plausible-sounding but ungrounded factual claims, a problem that worsens in multi-turn dialogue as context grows and early errors cascade. We introduce $\textbf{HalluHard}$, a challenging multi-turn hallucination benchmark with 950 seed questions spanning four high-stakes domains: legal cases, research questions, medical guidelines, and coding. We operationalize groundedness by requiring inline citations for factual assertions. To support reliable evaluation in open-ended settings, we propose a judging pipeline that iteratively retrieves evidence via web search. It can fetch, filter, and parse full-text sources (including PDFs) to assess whether cited material actually supports the generated content. Across a diverse set of frontier proprietary and open-weight models, hallucinations remain substantial even with web search ($\approx 30\%$ for the strongest configuration, Opus-4.5 with web search), with content-grounding errors persisting at high rates. Finally, we show that hallucination behavior is shaped by model capacity, turn position, effective reasoning, and the type of knowledge required.

幻觉评测多轮对话事实核查LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。