arXiv:2412.17259cs.CLcs.IR2024-12ACL被引 78

构建中文法律领域LLM Agent评估基准,覆盖真实司法场景。

LegalAgentBench: Evaluating LLM Agents in Legal Domain

  • 设计17个真实法律场景数据集与37种外部工具交互方式。
  • 标注300个任务,涵盖多跳推理与文书撰写,难度分级。
  • 引入中间步骤关键词分析,实现过程级进度评估。

随着大语言模型代理(LLM Agents)智能与自主性的提升,其在法律领域的应用潜力日益显现。然而,现有通用领域基准无法充分反映真实司法认知与决策的复杂性与细微差别。为此,我们提出LegalAgentBench,一个专为中文法律领域设计的综合性评估基准。该基准包含17个来自真实法律场景的语料库,提供37种与外部知识交互的工具。我们设计了可扩展的任务构建框架,并精心标注了300个任务,覆盖多跳推理、文书写作等多种类型,且具有不同难度层级,有效体现真实法律场景的复杂性。此外,除最终成功率外,LegalAgentBench还通过分析中间过程中的关键词,计算进展率,实现更细粒度的评估。我们对8个主流LLM进行了评估,揭示了现有模型在能力、局限及改进方向上的特征。LegalAgentBench为大模型在法律领域的实际应用设立了新标准,代码与数据已开源于\url{https://github.com/CSHaitao/LegalAgentBench}。

原文摘要 · Abstract (English)

With the increasing intelligence and autonomy of LLM agents, their potential applications in the legal domain are becoming increasingly apparent. However, existing general-domain benchmarks cannot fully capture the complexity and subtle nuances of real-world judicial cognition and decision-making. Therefore, we propose LegalAgentBench, a comprehensive benchmark specifically designed to evaluate LLM Agents in the Chinese legal domain. LegalAgentBench includes 17 corpora from real-world legal scenarios and provides 37 tools for interacting with external knowledge. We designed a scalable task construction framework and carefully annotated 300 tasks. These tasks span various types, including multi-hop reasoning and writing, and range across different difficulty levels, effectively reflecting the complexity of real-world legal scenarios. Moreover, beyond evaluating final success, LegalAgentBench incorporates keyword analysis during intermediate processes to calculate progress rates, enabling more fine-grained evaluation. We evaluated eight popular LLMs, highlighting the strengths, limitations, and potential areas for improvement of existing models and methods. LegalAgentBench sets a new benchmark for the practical application of LLMs in the legal domain, with its code and data available at \url{https://github.com/CSHaitao/LegalAgentBench}.

法律AILLM评估智能代理中文基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。