arXiv:2512.04578cs.CL2025-12ACL被引 3

构建中文法律通用智能评估基准,系统检验大模型法律能力。

LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence

  • 按维度-任务-能力框架设计,覆盖7维度11任务20项能力。
  • 基于真实案例和考题生成多选题,经人工与模型双重审核控漏。
  • 发现顶尖模型仍远低于人类律师水平,适合法律AI研发者使用。

法律通用智能(Legal GI)指人工智能在法律理解、推理与决策方面的综合能力,模拟法律专家跨领域的专业水平。现有评测体系偏重结果而缺乏对法律智能的系统性评估,制约了法律GI的发展。为此,我们提出LexGenius——一个面向大语言模型法律通用智能的专家级中文基准。该基准采用维度-任务-能力框架,涵盖7个维度、11项任务和20种能力。通过最新真实案例与考试题目构建选择题,结合人工与大模型双重审查,降低数据泄露风险,经多轮校验确保准确性与可靠性。我们对12个先进大模型进行了评估并深入分析,发现各模型在法律智能能力上存在显著差异,即使表现最优的模型也未达到人类法律专业人士水平。我们认为LexGenius可有效评估大模型的法律智能,推动法律通用智能发展。项目代码已开源:https://github.com/QwenQKing/LexGenius。

原文摘要 · Abstract (English)

Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulating the expertise of legal experts across domains. However, existing benchmarks are result-oriented and fail to systematically evaluate the legal intelligence of large language models (LLMs), hindering the development of legal GI. To address this, we propose LexGenius, an expert-level Chinese legal benchmark for evaluating legal GI in LLMs. It follows a Dimension-Task-Ability framework, covering seven dimensions, eleven tasks, and twenty abilities. We use the recent legal cases and exam questions to create multiple-choice questions with a combination of manual and LLM reviews to reduce data leakage risks, ensuring accuracy and reliability through multiple rounds of checks. We evaluate 12 state-of-the-art LLMs using LexGenius and conduct an in-depth analysis. We find significant disparities across legal intelligence abilities for LLMs, with even the best LLMs lagging behind human legal professionals. We believe LexGenius can assess the legal intelligence abilities of LLMs and enhance legal GI development. Our project is available at https://github.com/QwenQKing/LexGenius.

法律AI大模型评测通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。