构建加拿大判例法问答基准,评估法律RAG系统真实表现
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

- 基于真实法律查询和专家标注答案构建评测基准
- 8%-29%生成答案内容无法由检索文档支持,存在幻觉
- 开源嵌入模型性能可比闭源模型,自动评估有局限
基于RAG的法律助手机器人日益流行,但大模型幻觉问题仍威胁司法公正。现有评测多依赖合成查询,且缺乏对加拿大法律的覆盖。为此,我们提出CanLegalRAGBench,一个基于真实法律查询和判例法支撑的专家标注问答基准。评估显示,检索性能受设计选择影响显著;开源嵌入模型表现可媲美闭源模型。但自动评估会因系统召回相关替代文档而扣分。此外,生成答案常偏离标准答案,存在幻觉或过度细节,8%-29%的主张未获检索文档支持。该基准旨在推动法律RAG系统的持续改进。
原文摘要 · Abstract (English)
RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmarks have been developed to evaluate progress, many rely on synthetic queries rather than realistic legal scenarios. Moreover, Canadian law remains underrepresented in existing evaluations. To address this gap, we introduce CanLegalRAGBench, a Canadian legal QA benchmark based on realistic queries and expert-annotated answers grounded in case law. Our evaluation shows that retrieval performance is sensitive to design choices and that open-source embedding models are competitive with closed source models. However, it also reveals the limitation of automatic evaluations that penalize systems for retrieving alternative relevant documents. We also find that generated answers often diverge from gold responses, either with hallucinations or by producing overly detailed or irrelevant content, with 8-29% of claims not being supported by the retrieved documents. We hope this benchmark will help drive continued progress in addressing limitations of legal RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。