arXiv:2509.24922cs.AIcs.CL2025-09被引 3

构建面向法律推理的多智能体评估基准,验证大模型协作能力

MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning

  • 基于欧盟隐私法设计多角色协作框架,支持任务分解与专业化分工
  • 实验表明当前大模型在复杂法律推理中存在逻辑断裂与角色协同缺陷
  • 适合研究法律AI、多智能体系统与生成式模型协同的学者使用

多智能体系统(MAS)借助大语言模型(LLM)的强大能力,在处理复杂任务方面展现出巨大潜力。将MAS与法律任务结合是关键进展。尽管已有法律类LLM评测基准,但均未针对MAS的独特优势(如任务分解、角色专精、灵活训练)进行设计。评估方法的缺失制约了MAS在法律领域的应用。为此,我们提出MASLegalBench,一个专为多智能体系统设计的法律推理基准,采用演绎推理范式,以GDPR为应用场景,涵盖广泛背景知识和复杂推理流程,真实反映现实法律情境。我们人工设计多种角色型多智能体架构,并在多个先进LLM上进行大量实验。结果揭示了现有模型与架构在法律推理中的优劣及改进空间。

原文摘要 · Abstract (English)

Multi-agent systems (MAS), leveraging the remarkable capabilities of Large Language Models (LLMs), show great potential in addressing complex tasks. In this context, integrating MAS with legal tasks is a crucial step. While previous studies have developed legal benchmarks for LLM agents, none are specifically designed to consider the unique advantages of MAS, such as task decomposition, agent specialization, and flexible training. In fact, the lack of evaluation methods limits the potential of MAS in the legal domain. To address this gap, we propose MASLegalBench, a legal benchmark tailored for MAS and designed with a deductive reasoning approach. Our benchmark uses GDPR as the application scenario, encompassing extensive background knowledge and covering complex reasoning processes that effectively reflect the intricacies of real-world legal situations. Furthermore, we manually design various role-based MAS and conduct extensive experiments using different state-of-the-art LLMs. Our results highlight the strengths, limitations, and potential areas for improvement of existing models and MAS architectures.

法律AI多智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。