提出新框架让AI在证据不足时主动不判断,避免法律决策中的盲目自信。
Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication

- 设计结构化提示,强制AI先识别缺失信息再做判断。
- 新方法在证据不足时准确率提升至89%,且不乱下结论。
- 适合需要谨慎决策的法律、保险等高风险场景使用。
AI系统常因信息不足而过度自信地给出答案,这在法律领域尤为危险。本文聚焦失业保险裁定这一重要场景,与科罗拉多州劳工部合作,构建了首个系统性变化信息完整性的基准。评估发现,主流RAG方法在信息不足时平均准确率仅15%;高级提示法虽改善了模糊案例表现,却过度保守,连明确案件也拒绝判断。为此,我们提出结构化提示框架SPEC(Structured Prompting for Evidence Checklists),要求模型在决策前显式列出缺失证据。实验表明,SPEC实现89%的整体准确率,并在证据不足时正确推迟判断,证明法律AI的盲目自信是可被系统性解决的,而这种能力正是可靠辅助人类决策的前提。
原文摘要 · Abstract (English)
A well-known limitation of AI systems is presumptuousness: the tendency of AI systems to provide confident answers when information may be lacking. This challenge is particularly acute in legal applications, where a core task for attorneys, judges, and administrators is to determine whether evidence is sufficient to reach a conclusion. We study this problem in the important setting of unemployment insurance adjudication, which has seen rapid integration of AI systems and where the question of additional fact-finding poses the most significant bottleneck for a system that affects millions of applicants annually. First, through a collaboration with the Colorado Department of Labor and Employment, we secure rare access to official training materials and guidance to design a novel benchmark that systematically varies in information completeness. Second, we evaluate four leading AI platforms and show that standard RAG-based approaches achieve an average of only 15% accuracy when information is insufficient. Third, advanced prompting methods improve accuracy on inconclusive cases but over-correct, withholding decisions even on clear cases. Fourth, we introduce a structured framework requiring explicit identification of missing information before any determination (SPEC, Structured Prompting for Evidence Checklists). SPEC achieves 89% overall accuracy, while appropriately deferring when evidence is insufficient -- demonstrating that presumptuousness in legal AI is systematic but addressable, and that doing so is a necessary step towards systems that reliably support, rather than supplant, human judgment wherever decisions must await sufficient evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。