arXiv:2509.21208cs.CL2025-09EMNLP被引 3

构建中文法律大模型测评基准,揭示现有模型法律知识严重不足

CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis

  • 构建306部中国法律的细粒度语料库,精确到条款层级并标注修订时间
  • 测试254个最高法案例推理任务,发现多数模型无法准确引用法律条文
  • 强调法律推理需结合精准检索与强逻辑能力,适合法律AI研究者参考

大型语言模型在处理法律文本和引用法规方面日益重要,但其可靠性常因通用预训练中混杂法律文本而受损。本文提出CLaw,一个专为评估中文法律知识与推理能力设计的新基准。该基准包含两个核心部分:(1) 全面、细粒度的语料库,涵盖全部306部中国国家法律,按子条款级别分割,并标注精确的历史修订时间,共64,849个条目;(2) 254个基于中国最高人民法院材料的案例推理题,用于评估法律知识的实际应用。实证评估显示,多数主流大模型在忠实复现法律条文方面表现不佳。由于准确检索与引用法律条文是法律推理的基础,这一缺陷严重削弱了其回答可信度。我们主张,实现可信赖的法律推理需要精准知识检索(如监督微调或检索增强生成)与强大泛化推理能力的协同。本工作为提升特定领域大模型推理能力,尤其是在复杂法律场景中,提供了关键基准与深刻洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly tasked with analyzing legal texts and citing relevant statutes, yet their reliability is often compromised by general pre-training that ingests legal texts without specialized focus, obscuring the true depth of their legal knowledge. This paper introduces CLaw, a novel benchmark specifically engineered to meticulously evaluate LLMs on Chinese legal knowledge and its application in reasoning. CLaw comprises two key components: (1) a comprehensive, fine-grained corpus of all 306 Chinese national statutes, segmented to the subparagraph level and incorporating precise historical revision timesteps for rigorous recall evaluation (64,849 entries), and (2) a challenging set of 254 case-based reasoning instances derived from China Supreme Court curated materials to assess the practical application of legal knowledge. Our empirical evaluation reveals that most contemporary LLMs significantly struggle to faithfully reproduce legal provisions. As accurate retrieval and citation of legal provisions form the basis of legal reasoning, this deficiency critically undermines the reliability of their responses. We contend that achieving trustworthy legal reasoning in LLMs requires a robust synergy of accurate knowledge retrieval--potentially enhanced through supervised fine-tuning (SFT) or retrieval-augmented generation (RAG)--and strong general reasoning capabilities. This work provides an essential benchmark and critical insights for advancing domain-specific LLM reasoning, particularly within the complex legal sphere.

法律AI大模型评测知识检索推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。