arXiv:2605.21071cs.CLcs.AI2026-05被引 2

构建法律领域细粒度RAG评估基准,支持多语言与多用户场景。

Fine-grained Claim-level RAG Benchmark for Law

论文配图:Fine-grained Claim-level RAG Benchmark for Law
图 1 · 摘自论文原文
  • 提出ClaimRAG-LAW数据集,覆盖专家与非专家需求。
  • 发现现有法律RAG在检索与生成上仍存在显著幻觉。
  • 支持中英双语,适用于法律AI系统精细化评估。

大语言模型的快速发展正推动语义搜索向问答范式转变,用户提问后由LLM生成回答。在法律等高风险领域,检索增强生成(RAG)常被用于缓解生成内容的幻觉问题。然而,已有研究表明,无论通用或法律专用的RAG系统仍以不同比率产生幻觉,因此需要细粒度评估。现有法律RAG评估框架缺乏对检索与生成性能的分别分析能力,且主要局限于英语和法律专家查询,忽视了非专家需求。本文提出ClaimRAG-LAW,一个涵盖中英双语、面向专家与非专家、包含多样化真实场景问题类型的综合性法律RAG评估数据集。进一步应用细粒度评估框架对主流法律RAG系统进行测试,揭示了其在检索、生成及主张级分析方面的局限性。

原文摘要 · Abstract (English)

The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses. In high-stake domains such as law, retrieval-augmented generation (RAG) is commonly used to mitigate hallucinations in generated responses. Nonetheless, prior work shows that RAG systems, whether general-purpose or legal-specific, still hallucinate at varying rates, making fine-grained evaluation essential. Despite the need, existing evaluation frameworks for legal RAG systems lack the granularity required to provide detailed analysis of retrieval and generation performance separately. Moreover, current benchmarks are largely English-only and centered on legal expert queries, overlooking non-expert needs. We introduce ClaimRAG-LAW, a comprehensive dataset for legal RAG that supports French and English, targets both experts and non-experts, and includes diverse question types reflecting realistic scenarios. We further apply a fine-grained evaluation framework of state-of-the-art legal RAG systems, revealing limitations in retrieval, generation, and claim-level analysis in the legal domain.

法律AIRAG评估多语言细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。