arXiv:2510.17463cs.AI2025-10中稿 · presentation as a …

法律判决因程序干预存在不确定性,影响机器学习模型效果。

Label Indeterminacy in AI & Law

  • 用历史案件结果训练模型时,忽略程序干预导致标签不明确
  • 不同标签构建方式使模型表现差异显著,影响预测可靠性
  • 适合关注AI司法应用可信性与数据偏差的研究者

机器学习在法律领域的应用日益广泛,通常以过往案件判决结果作为训练目标。然而,法律判决常受和解、上诉等人为程序因素影响,这些过程未被大多数机器学习方法捕捉,导致标签具有不确定性:若无这些干预,结果可能完全不同。我们指出,法律类机器学习必须考虑标签不确定性。现有方法可推断此类不确定标签,但均基于无法验证的假设。以欧洲人权法院案件分类为例,我们证明标签构建方式会显著影响模型行为。因此,标签不确定性是AI与法律交叉领域的重要议题,其对模型表现具有实质性影响。

原文摘要 · Abstract (English)

Machine learning is increasingly used in the legal domain, where it typically operates retrospectively by treating past case outcomes as ground truth. However, legal outcomes are often shaped by human interventions that are not captured in most machine learning approaches. A final decision may result from a settlement, an appeal, or other procedural actions. This creates label indeterminacy: the outcome could have been different if the intervention had or had not taken place. We argue that legal machine learning applications need to account for label indeterminacy. Methods exist that can impute these indeterminate labels, but they are all grounded in unverifiable assumptions. In the context of classifying cases from the European Court of Human Rights, we show that the way that labels are constructed during training can significantly affect model behaviour. We therefore position label indeterminacy as a relevant concern in AI & Law and demonstrate how it can shape model behaviour.

AI与法律标签不确定性司法人工智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。