arXiv:2608.01559cs.AIcs.CL2026-08

测试对抗自对弈能否提升法律推理,结果发现竞争无显著效果。

Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result

  • 设计可验证的生存奖励机制,确保引用真实可靠
  • 四类测试均显示竞争组无明显优势,胜率约49%-50%
  • 揭示虚假指标陷阱,强调环境可验证性比竞争更重要

对抗自对弈在法律推理中具有吸引力:让学生模型起草论点,由对手攻击,若论点经受住攻击则给予奖励。我们构建了基于可验证引用的‘生存’奖励机制,通过引用验证器检查双方引文,确保判断基于事实而非修辞,且伪造引用自动失效。我们提出一个关键问题:竞争成分(对手与生存奖励)是否在非竞争训练基础上带来实际提升?通过四项独立测试——自举对比、双种子复现、逐案对抗鲁棒性对比、盲评生成论点对决,以及一次强化对手的预实验,结果均未显示可靠优势。盲评胜率49%(二项分布p≈1.000),强化对手实验胜率50%(32:32,p≈1.000)。早期看似+29%的优势实为小样本偏差。本文报告这一真实负结果,强调其可复现性与潜在陷阱:初始有前景的指标随数据量增加反转,对抗鲁棒性指标在对手不再引用标准答案时悄然退化为普通召回率。该结论与配套编码领域研究一致,表明多教师课程的价值源于可验证环境,而非竞争本身。

原文摘要 · Abstract (English)

Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the student when its argument survives the attack. We designed exactly such a training signal -- a verifiable "survival" reward in which both the student's cited authorities and the adversary's counter-authorities are checked by a citation verifier, so that survival is decided on verified grounds rather than rhetoric, and fabricated citations are automatically neutralized. We then asked a narrow but important question: does the competitive component itself -- the adversary and the survival reward -- add anything on top of an otherwise identical non-competitive training run? Across four independent tests -- a bootstrap comparison, a two-seed replication, a paired per-case adversarial-robustness comparison, and a blinded head-to-head judgment of generated arguments, plus a follow-up pilot with a deliberately strengthened self-play adversary -- the competitive component produced no reliable benefit. The blinded judgment gave a 49% win rate (binomial p approx. 1.000); the strengthened-adversary pilot gave a 50% win rate (32:32, p approx. 1.000). An early apparent +29% advantage reversed and proved to be a small-sample artifact. We report this as an honest negative result. The value of the paper is reproducibility and the sharing of concrete pitfalls: an initially promising metric that inverted on more data, and an adversarial-robustness metric that silently collapsed to plain recall once the adversary stopped citing the same authorities as the gold answer. This null is consistent with, and reconfirms in the legal domain, the conclusion of the companion coding-domain study (Kim, 2026, arXiv:2607.08255) that the value of multi-teacher curricula arises from constructing a verifiable environment rather than from competition itself.

法律推理对抗训练负结果

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。