arXiv:2608.14210cs.CLcs.AI2026-08

分析八种法律RAG系统幻觉问题,发现最差情况近半答案不实。

How Much Do Legal RAG Systems Still Hallucinate?

  • 基于条款级与答案级评估,量化法律RAG幻觉密度与严重性。
  • 最佳系统幻觉率低于10%,最差达48%,且虚假前提问题幻觉更严重。
  • 适用于法律AI评测、司法辅助系统开发者及合规研究者。

幻觉是法律领域检索增强生成(RAG)系统的主要挑战,因无根据的回答可能导致严重后果。为深入理解该问题,我们在两个法律语料库(GDPR英文版与法国全国民事法典)上对八种法律RAG系统进行细粒度分析。采用条款级与答案级评估,报告幻觉密度与严重性,分析不同问题类型与用户角色下的表现,并在142个由法律专家撰写的独立问题集上验证结果。结果显示,幻觉仍普遍存在,最佳系统幻觉率低于10%,最差系统高达48%。此外,包含错误假设的虚假前提问题在人工设计的问题中引发极高幻觉率。

原文摘要 · Abstract (English)

Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR (in English) and a national civil law (in French). Using claim-level and answer-level evaluation, we report on hallucination density and severity, analyze performance across question categories and user personas, and validate our findings on an independent set of 142 legal-expert-authored questions. Our results show that hallucinations remain pervasive, ranging from less than 10% of responses for the best-performing systems to nearly half in the worst case. We further find that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.

法律AIRAG幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。