arXiv:2510.03405cs.MAcs.AI2025-10中稿 · EMNLP被引 2

用多智能体模拟法律程序,发现系统性漏洞。

LegalSim: Multi-Agent Simulation of Legal Systems for Discovering Procedural Exploits

  • 构建可配置的法律对抗仿真环境,智能体按规则行动
  • 发现成本膨胀、时间压力等合法但有害的策略链
  • 适合法律AI安全测试与规则体系红队演练

我们提出LegalSim,一个模块化的多智能体对抗法律程序仿真系统,用于探索人工智能如何利用编码化规则中的程序漏洞。原告与被告智能体在受限动作空间(如证据开示请求、动议、会面协商、制裁)中选择行为,由具备校准通过率、成本分摊和制裁倾向的随机法官模型裁决结果。对比四种策略:PPO、基于LLM的上下文老虎机、直接LLM策略及人工设计启发式规则。评估不以二元胜负为目标,而是使用有效胜率和综合漏洞得分(包含对手成本抬升、日程压力、低理据下施压和规则合规余量)。在可配置场景(如破产暂停、双边复审、税务程序)与异质法官设置下,观察到涌现的“漏洞链”,如成本膨胀的证据开示序列与合法但系统有害的时间压力策略。跨对弈与Bradley-Terry评分显示,PPO胜率最高,老虎机在多数对手中表现最稳定,LLM次之,启发式最弱。结果在不同法官设定下保持稳健,揭示了系统性漏洞链,推动对法律规则体系进行红队测试,而不仅是模型层面验证。

原文摘要 · Abstract (English)

We present LegalSim, a modular multi-agent simulation of adversarial legal proceedings that explores how AI systems can exploit procedural weaknesses in codified rules. Plaintiff and defendant agents choose from a constrained action space (for example, discovery requests, motions, meet-and-confer, sanctions) governed by a JSON rules engine, while a stochastic judge model with calibrated grant rates, cost allocations, and sanction tendencies resolves outcomes. We compare four policies: PPO, a contextual bandit with an LLM, a direct LLM policy, and a hand-crafted heuristic; Instead of optimizing binary case outcomes, agents are trained and evaluated using effective win rate and a composite exploit score that combines opponent-cost inflation, calendar pressure, settlement pressure at low merit, and a rule-compliance margin. Across configurable regimes (e.g., bankruptcy stays, inter partes review, tax procedures) and heterogeneous judges, we observe emergent ``exploit chains'', such as cost-inflating discovery sequences and calendar-pressure tactics that remain procedurally valid yet systemically harmful. Evaluation via cross-play and Bradley-Terry ratings shows, PPO wins more often, the bandit is the most consistently competitive across opponents, the LLM trails them, and the heuristic is weakest. The results are stable in judge settings, and the simulation reveals emergent exploit chains, motivating red-teaming of legal rule systems in addition to model-level testing.

法律AI多智能体红队测试规则漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。