首个动态法律环境评测框架,揭示大模型在真实法律场景中的执行短板。
Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
- 构建动态法律环境J1-ENVS,模拟中国法律实践的六类场景。
- 17个大模型在动态环境中平均表现不足60%,连GPT-4o也未达标。
- 专设评估体系,兼顾任务完成与程序合规性,适合法律AI研究者参考。
静态基准与真实法律实践的动态性之间存在显著鸿沟,制约了法律智能的发展。为此,我们提出J1-ENVS,首个面向基于大语言模型(LLM)智能体的交互式、动态法律环境。该环境由法律专家指导,涵盖中国法律实践中三类复杂度的六种典型场景。我们进一步设计了J1-EVAL,一个细粒度评估框架,用于衡量不同法律能力水平下的任务表现与程序合规性。对17个LLM智能体的广泛实验表明,尽管多数模型具备扎实的法律知识,但在动态环境中仍难以有效执行程序流程。即使是最先进的GPT-4o,整体表现也未超过60%。这些发现凸显了实现动态法律智能的持续挑战,并为未来研究提供了关键洞见。
原文摘要 · Abstract (English)
The gap between static benchmarks and the dynamic nature of real-world legal practice poses a key barrier to advancing legal intelligence. To this end, we introduce J1-ENVS, the first interactive and dynamic legal environment tailored for LLM-based agents. Guided by legal experts, it comprises six representative scenarios from Chinese legal practices across three levels of environmental complexity. We further introduce J1-EVAL, a fine-grained evaluation framework, designed to assess both task performance and procedural compliance across varying levels of legal proficiency. Extensive experiments on 17 LLM agents reveal that, while many models demonstrate solid legal knowledge, they struggle with procedural execution in dynamic settings. Even the SOTA model, GPT-4o, falls short of 60% overall performance. These findings highlight persistent challenges in achieving dynamic legal intelligence and offer valuable insights to guide future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。