arXiv:2409.20288cs.CL2024-09NeurIPS被引 41

构建首个大规模中文法律大模型评测基准,全面评估法律认知与伦理能力。

LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models

  • 基于法律认知能力新分类,系统组织23类法律任务
  • 涵盖14,150道题,是目前最大的中文法律评测数据集
  • 适合法律AI研发、评测人员及政策制定者参考

大语言模型在自然语言处理任务中取得显著进展,在法律领域也展现出巨大潜力。然而,法律应用对准确性、可靠性与公平性要求极高,若未充分评估大模型的潜在局限性便直接应用于法律系统,可能带来重大风险。为此,我们提出一个标准化的综合性中文法律评测基准LexEval。该基准具有三大特点:(1)能力建模:提出新的法律认知能力分类体系,用于组织不同任务;(2)规模:据我们所知,LexEval是当前最大的中文法律评估数据集,包含23个任务和14,150个问题;(3)数据:融合格式化现有数据集、考试数据集及法律专家新标注数据,全面评估大模型的各项能力。该基准不仅关注模型对基础法律知识的应用能力,还重点考察其应用中的伦理问题。我们评估了38个开源与商业大模型,获得若干有趣发现。实验结果为发展中文法律系统及大模型评测流程提供了重要启示。LexEval数据集与排行榜已公开于https://github.com/CSHaitao/LexEval,将持续更新。

原文摘要 · Abstract (English)

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, reliability, and fairness. Applying existing LLMs to legal systems without careful evaluation of their potential and limitations could pose significant risks in legal practice. To this end, we introduce a standardized comprehensive Chinese legal benchmark LexEval. This benchmark is notable in the following three aspects: (1) Ability Modeling: We propose a new taxonomy of legal cognitive abilities to organize different tasks. (2) Scale: To our knowledge, LexEval is currently the largest Chinese legal evaluation dataset, comprising 23 tasks and 14,150 questions. (3) Data: we utilize formatted existing datasets, exam datasets and newly annotated datasets by legal experts to comprehensively evaluate the various capabilities of LLMs. LexEval not only focuses on the ability of LLMs to apply fundamental legal knowledge but also dedicates efforts to examining the ethical issues involved in their application. We evaluated 38 open-source and commercial LLMs and obtained some interesting findings. The experiments and findings offer valuable insights into the challenges and potential solutions for developing Chinese legal systems and LLM evaluation pipelines. The LexEval dataset and leaderboard are publicly available at \url{https://github.com/CSHaitao/LexEval} and will be continuously updated.

法律AI大模型评测中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。