arXiv:2506.03913cs.CLcs.LG2025-06被引 2

用机器学习评估难民裁决公平性,发现方法不一致且难反映法律实质。

When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning

  • 对比三种机器学习方法在真实裁决数据上的表现
  • 预测模型依赖程序特征而非法律内容,语义聚类无法捕捉法律推理
  • 提醒法律公平需结合法律逻辑与制度背景,非纯数据可解

随着机器学习在法律决策公平性评估中的应用日益广泛,特别是在难民裁决等高风险领域,其有效性仍存疑。本文基于包含59,000+份加拿大难民裁决的真实数据集AsyLex,实证评估了三种常见机器学习方法:基于特征的分析、语义聚类和预测建模。实验显示,这些方法产生分歧甚至矛盾的结果;预测模型常依赖程序性与上下文特征,而非法律实质内容;语义聚类无法有效捕捉实质性法律推理。研究揭示统计公平评估在法律裁量领域存在根本局限,挑战了“统计规律即公平”的假设,强调法律公平评估需结合法律推理与制度背景,现有计算方法难以胜任。

原文摘要 · Abstract (English)

Legal decisions are increasingly evaluated for fairness, consistency, and bias using machine learning (ML) techniques. In high-stakes domains like refugee adjudication, such methods are often applied to detect disparities in outcomes. Yet it remains unclear whether statistical methods can meaningfully assess fairness in legal contexts shaped by discretion, normative complexity, and limited ground truth. In this paper, we empirically evaluate three common ML approaches (feature-based analysis, semantic clustering, and predictive modeling) on a large, real-world dataset of 59,000+ Canadian refugee decisions (AsyLex). Our experiments show that these methods produce divergent and sometimes contradictory signals, that predictive modeling often depends on contextual and procedural features rather than legal features, and that semantic clustering fails to capture substantive legal reasoning. We show limitations of statistical fairness evaluation, challenge the assumption that statistical regularity equates to fairness, and argue that current computational approaches fall short of evaluating fairness in legally discretionary domains. We argue that evaluating fairness in law requires methods grounded not only in data, but in legal reasoning and institutional context.

法律人工智能公平性评估难民裁决机器学习局限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。