测试不同协作模式在法律推理中的表现,发现人多未必更好。
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

- 给多个智能体分配法律专家角色,通过辩论提升推理能力
- 相比单个智能体,准确率最高提升8%,但辩论轮次过多会出错
- 适合研究法律AI协作机制的学者与开发者参考
尽管多智能体辩论(MAD)框架在一般推理中展现出巨大潜力,但在高度结构化、知识密集的法律领域,其有效性仍缺乏系统研究。本文提出法律多智能体辩论(L-MAD)框架,系统评估法律文本蕴含任务中不同的辩论结构与聚合方法。通过为多个智能体分配不同专家角色,L-MAD相较强基线模型准确率最高提升8%。分析辩论规模的影响发现:增加智能体数量可降低不一致性和提高准确性,但延长讨论轮次会导致负面的“过度审议漂移”——智能体相互强化错误。研究结果明确了高风险法律推理环境中协作式多智能体系统的实用边界与安全范围。
原文摘要 · Abstract (English)
While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains under-explored. In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within Legal Textual Entailment. By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8\%. Furthermore, analyzing how debate scales reveals a clear trade-off: increasing the agent population reduces inconsistency and improves accuracy, whereas extending discussion rounds induces a detrimental \textit{over-deliberation drift} where agents reinforce each other's mistakes. Ultimately, our findings outline the practical boundaries and safety margins of deploying collaborative multi-agent systems in high-stakes legal reasoning environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。