arXiv:2601.12419cs.CL2026-01

对比法律专家与AI解释,发现模型理由与专家不一致。

Legal Experts Disagree With Rationale Extraction Techniques for Explaining ECtHR Case Outcome Classification

  • 用新构建的欧洲人权法院数据集,对比两种可解释性方法
  • 模型解释虽数学上准确,但法律合理性遭专家质疑
  • 适合关注AI司法解释可信度的研究者和法律科技从业者

大语言模型在法律领域应用需具备可解释性以保障信任与透明。本文聚焦欧洲人权法院(ECtHR)判决结果预测任务,构建了经精心标注的正负样本数据集。现有研究提出特定任务与通用解释方法,但尚不清楚哪种更优。为此,本文提出一种对比分析框架,重点评估两种基于文本片段的理性提取技术。通过归一化充分性与全面性衡量忠实度,结合法律专家对提取理由的判断评估合理性。同时测试大模型作为裁判者的可行性,以专家评价为基准。实验表明,尽管模型解释在数学指标上表现良好,其生成理由与法律专家认知存在显著差异。代码已公开于 https://github.com/trusthlt/IntEval。

原文摘要 · Abstract (English)

Interpretability is critical for applications of large language models (LLMs) in the legal domain, where trust and transparency are essential. A central NLP task in this setting is legal outcome prediction, where models forecast whether a court will find a violation of a given right. We study this task on decisions from the European Court of Human Rights (ECtHR), introducing a new ECtHR dataset with carefully curated positive (violation) and negative (non-violation) cases. Existing works propose both task-specific approaches and model-agnostic techniques to explain downstream performance, but it remains unclear which techniques best explain legal outcome prediction. To address this, we propose a comparative analysis framework for model-agnostic interpretability methods. We focus on two rationale extraction techniques that justify model outputs with concise, human-interpretable text fragments from the input. We evaluate faithfulness via normalized sufficiency and comprehensiveness metrics, and plausibility via legal expert judgments of the extracted rationales. We also assess the feasibility of using LLM-as-a-Judge, using these expert evaluations as reference. Our experiments on the new ECtHR dataset show that models' "reasons" for predicting violations differ substantially from those of legal experts, despite strong faithfulness scores. The source code of our experiments is publicly available at https://github.com/trusthlt/IntEval.

可解释AI法律AI模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。