arXiv:2605.26530cs.AI2026-05

让法律大模型只对关键变更敏感,提升可信度

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

论文配图:Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning
图 1 · 摘自论文原文
  • 构建可区分法律相关/无关变化的评估体系
  • 现有模型对无关细节过度敏感,易混淆相似法条
  • 提出基于形式推理的对抗性框架LexGuard,增强可靠性

法律推理需区分重要变更与无关变动。法律AI应在法律无关扰动下保持稳定,但在涉及法律实质内容的变更时应作出响应。本文将此要求建模为法律相关性敏感评估问题:大模型仅应对法律相关变更敏感。我们提出一个统一评估套件,覆盖司法公平、鲁棒性与法条混淆场景下的应变与不应变测试。实验表明,现有法律大模型系统性地对法律无关变化敏感,且难以区分相关法律要素与法条规则。为此,我们提出LexGuard——一种基于形式推理的对抗多智能体框架。该框架将法规形式化为可执行约束,利用对抗智能体提取竞争性事实-法条论证,并调用SMT求解器验证法律满足性与逻辑一致性。实验显示,LexGuard能有效降低对操纵性表述的脆弱性,提升相似法条间的歧义化解能力,减弱无关属性影响,并增强良性重述下的稳定性。结果表明,法律可信度不仅依赖准确率,更需对法律实质变更具备校准后的敏感性。

原文摘要 · Abstract (English)

Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations, but should change when perturbations alter legally material points. We formulate this requirement as a legal-relevance-sensitive evaluation problem: LLMs should only be sensitive to the legally relevant change. We introduce a unified evaluation suite covering should-change and should-not-change evaluation across judicial fairness, robustness, and statute-confusion scenarios. Our evaluation shows that existing legal LLMs are systematically sensitive to legally irrelevant variations and often fail to distinguish related legal elements and statutory rules. To mitigate these failures, we present LexGuard, an adversarial multi-agent framework grounded in formal reasoning. LexGuard formalizes statutes into executable constraints, uses adversarial agents to extract competing fact-statute arguments, and invokes SMT solvers to verify legal satisfaction and logical consistency. Experiments show that LexGuard improves legal reasoning reliability by reducing vulnerability to manipulative framing, improving disambiguation among similar statutes, limiting the influence of legally irrelevant attributes, and increasing consistency under benign reformulations. We show that legal trustworthiness requires not only accuracy, but calibrated sensitivity to legally material changes.

法律AI可信推理形式验证对抗评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。