提升法律判决预测的可靠性,让模型学会何时该不说话。
The Craft of Selective Prediction: Towards Reliable Case Outcome Classification -- An Empirical Study on European Court of Human Rights Cases
- 通过选择性预测框架,测试不同预训练数据、信心估计方法和损失函数的影响。
- 大模型易过度自信,蒙特卡洛丢弃法给出可靠信心估计,正则化可缓解过自信。
- 首个系统研究法律文本预测可信度的工作,适合关注司法AI可解释性的研究者。
在法律自然语言处理的高风险决策任务中,如案件结果分类(COC),量化模型的预测置信度至关重要。置信度估计有助于人类在模型把握不足或错误代价高昂时做出更明智的判断。然而,现有大多数COC研究更关注任务性能而非模型可靠性。本文针对欧洲人权法院(ECtHR)案例,对预训练语料库、置信度估计器和微调损失函数等设计选择如何影响选择性预测框架下的模型可靠性进行了实证研究。实验表明,多样且领域相关的预训练语料库有助于更好校准;大模型倾向于过度自信,蒙特卡洛丢弃法能产生可靠的置信度估计,而置信度误差正则化可有效缓解过自信问题。据我们所知,这是法律NLP中首次系统探索选择性预测的研究。研究强调了提升置信度度量与增强法律领域模型可信度的迫切需求。
原文摘要 · Abstract (English)
In high-stakes decision-making tasks within legal NLP, such as Case Outcome Classification (COC), quantifying a model's predictive confidence is crucial. Confidence estimation enables humans to make more informed decisions, particularly when the model's certainty is low, or where the consequences of a mistake are significant. However, most existing COC works prioritize high task performance over model reliability. This paper conducts an empirical investigation into how various design choices including pre-training corpus, confidence estimator and fine-tuning loss affect the reliability of COC models within the framework of selective prediction. Our experiments on the multi-label COC task, focusing on European Court of Human Rights (ECtHR) cases, highlight the importance of a diverse yet domain-specific pre-training corpus for better calibration. Additionally, we demonstrate that larger models tend to exhibit overconfidence, Monte Carlo dropout methods produce reliable confidence estimates, and confident error regularization effectively mitigates overconfidence. To our knowledge, this is the first systematic exploration of selective prediction in legal NLP. Our findings underscore the need for further research on enhancing confidence measurement and improving the trustworthiness of models in the legal domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。