arXiv:2606.23716cs.CYcs.AI2026-06中稿 · ICML

现有法律AI评测不真实,需测试普通人写诉状时模型的表现。

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice

论文配图:Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice
图 1 · 摘自论文原文
  • 用普通人原始诉状输入测试模型,而非专家整理的干净数据。
  • 实验证明模型在杂乱输入下性能大幅下降,与理想情况差距大。
  • 适合关注法律公平性、AI可及性的研究者和政策制定者。

法律AI基准测试常假设大语言模型能提升司法可及性,尤其帮助无法聘请律师的当事人理解并行使权利。然而,当前基准测试使用经过法律专家预处理的输入,仅衡量模型表现的上限。真正的司法可及性依赖于下限:模型在普通人(自诉人)提交的含噪声叙述、隐藏事实、遗漏信息、民间法律观念及表面错误的输入下的表现。这些干扰与通用机器学习中已知的模型退化现象类似,包括长上下文敏感性、表述不完整、幻觉和拼写扰动。本文结合自诉文献与机器学习研究,对法律基准LEXam进行小扰动实验,揭示了两种边界间的显著差距。若模型开发继续聚焦于仅测量上限的基准,此差距可能被掩盖甚至扩大。因此,呼吁建立直接评估模型在类自诉输入下鲁棒性的法律基准,使法律AI提升司法可及性的主张具备可检验性。

原文摘要 · Abstract (English)

Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights. We argue that current benchmarks are not equipped to support this assumption because they evaluate legal reasoning over inputs that have already been preprocessed by legal experts, which measures the upper bound of model performance. Access to justice depends on a lower bound: how models perform when inputs come from pro se litigants, whose prompts may contain noisy narratives, buried facts, omissions, folk-legal assumptions, and surface-level errors. These degradations are comparable to conditions under which LLMs are known to degrade in the general machine learning literature, including long-context sensitivity, underspecification, hallucination, and typographical perturbations. We connect evidence from pro se literature with this body of machine learning research and present a small perturbation experiment on LEXam, a legal benchmark, to illustrate the gap between these two bounds. If model development continues to focus on benchmarks that measure only the upper bound, this gap may remain hidden or even widen. We conclude by calling for legal benchmarks that directly measure robustness under pro se-like inputs so that access-to-justice claims about legal AI can become empirically testable.

法律AI司法可及模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。