arXiv:2607.04261cs.AI2026-07被引 1

发现法律判决模型依赖文本线索造假,去掉这些线索后仍能准确预测。

Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal

  • 用英国劳动法庭3.3万份案件文本测试模型,分析其是否依赖事后描述中的暗示信息
  • 仅用4%的泄露特征训练的模型就超越人类专家,说明性能被虚假线索夸大
  • 屏蔽泄露信息后模型性能几乎不变,证明仍有真正可学习的法律规律

当前法律判决预测(LJP)受限于对事后司法材料的依赖,导致模型更像在做回溯分类而非真实预测。本文通过研究英国劳动法庭(UKET)案件的申述结果预测,实证检验了其中的捷径学习现象。基于包含33,158个独立申诉的语料库,我们从申诉文本和LLM提取的案件摘要中预测结果,评估了从可解释的TF-IDF分类器到黑箱大模型等多种方法。尽管整体预测表现看似良好,但研究表明,这种性能主要源于源文本的回溯性质。按人工判断的泄漏程度分层测试数据发现,当结果提示信息嵌入叙述时,模型性能显著提升。此外,仅使用4%的泄露特征训练的模型即达到高精度,甚至超过人类专家。这证实了LJP性能可能因语言伪象而被夸大。然而,这一缺陷并非不可修复:将事后判决视为潜在污染文本并主动审计,通过遮蔽泄漏特征重训模型,仅导致宏平均F1轻微下降。因此,尽管模型会利用可用捷径,但在去除这些干扰后仍能捕捉有效的预测信号。

原文摘要 · Abstract (English)

Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLM-extracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1. Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.

法律人工智能捷径学习模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。