arXiv:2606.27069cs.CL2026-06中稿 · the AI for Law Wor…被引 1

区分案件事实与法官裁量,用门控多任务学习提升法律判决预测可解释性。

Towards Explainable Adjudicative Variance: Quantifying Judicial Discretion via Gated Multi-Task Learning

论文配图:Towards Explainable Adjudicative Variance: Quantifying Judicial Discretion via Gated Multi-Task Learning
图 1 · 摘自论文原文
  • 设计门控多任务架构,显式分离事实与法官背景的影响路径。
  • 在13,937份英国劳资仲裁判例上实现新基准,尤其提升模糊罕见类别的准确率。
  • 模型参数少10倍,且可解释性强,适合司法透明与政策评估场景。

法律判决预测需剥离客观案情与裁判语境的影响。基于事实的裁决依赖证据,而技术性处理可能受法官裁量影响。本文提出一种法官感知的门控多任务学习架构,显式建模这一差异。引入细粒度判决分类体系以监督编码器,施加结构正则化,实现语义路径解耦。该精细化法律课程使门控融合机制能动态调节对法官身份的依赖。在13,937份英国就业法庭判决数据集上评估,相比微调Gemma-4 26B-A4B主干模型(将法官身份和分类体系作为提示词或自回归输出),两组上下文信号在单一生成通道中仅弱关联。而采用LoRA适配的Gemma-4编码器与门控架构结合,显著超越生成式微调基线,参数量减少一个数量级,且在最模糊、最罕见的判决类别上表现最佳。模型不仅更准确,还具备可解释性:学习到的法官嵌入与校准曲线可定位裁量主导的案例。结果表明,在身份条件化法律判决分类中,条件接口的选择远比模型规模重要——可微分的结构化组合优于大模型上的提示工程。

原文摘要 · Abstract (English)

Legal outcome prediction must disentangle objective case facts from adjudicative context. Merit-based rulings rely on factual evidence while technical disposals may hinge on judicial discretion. We propose a Judge-Aware Gated Multi-Task Learning architecture that explicitly models this distinction. We introduce a fine-grained outcome taxonomy to supervise the encoder, enforcing a structural regularization that disentangles distinct semantic pathways. This granular legal curriculum enables our Gated Fusion mechanism to dynamically modulate reliance on judge identity. We evaluate our approach on 13,937 UK Employment Tribunal decisions. We benchmark our design against supervised fine-tuning (SFT) of a Gemma-4 26B-A4B backbone, in which judge identity and the taxonomy are injected as prompt tokens or autoregressive output targets. The two contextual signals compose only weakly when forced through a single autoregressive channel. In contrast, coupling a LoRA-adapted Gemma-4 encoder with our gated architecture defines a new state of the art on this benchmark while requiring an order of magnitude fewer trainable parameters than the generative SFT baselines, with gains concentrated on the most ambiguous and rarest outcome classes. Beyond accuracy, the architecture is interpretable; learned judge embeddings and calibration profiles localize the cases where adjudicative context drives the prediction. These results indicate that, for identity-conditioned classification of legal outcomes, the choice of conditioning interface dominates scale: differentiable structured composition yields more accurate, more parameter-efficient models than prompt-based composition over a substantially larger backbone.

法律AI可解释性多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。