arXiv:2601.22742cs.CL2026-01

构建法律判决错误检测基准,评估模型诊断能力

AR-BENCH: Benchmarking Legal Reasoning with Judgment Error Detection, Classification and Correction

  • 提出新任务APPELLATE REVIEW,聚焦判决后错误检测与修正
  • 构建含8700份标注判决的AR-BENCH数据集,支持细粒度分析
  • 14个大模型在法律适用错误识别上表现有限,揭示改进空间

法律判决因案件复杂性和法律概念抽象性可能包含错误,而现有上诉审查机制面临案量激增带来的效率压力。尽管当前法律AI研究集中于判决预测和法律文书生成,但判决审查任务在目标与范式上存在本质差异:其核心在于判决发布后的错误检测、分类与修正,属于异常检测而非预测或生成。为填补这一研究空白,我们提出新任务APPELLATE REVIEW,旨在评估模型在法律实践中的诊断推理与可靠性。同时构建了新型基准数据集AR-BENCH,包含8,700份精细标注的判决及34,617份补充语料。通过评估14个大语言模型,我们揭示了现有模型在识别法律适用错误方面的显著局限,为未来改进提供了实证依据。

原文摘要 · Abstract (English)

Legal judgments may contain errors due to the complexity of case circumstances and the abstract nature of legal concepts, while existing appellate review mechanisms face efficiency pressures from a surge in case volumes. Although current legal AI research focuses on tasks like judgment prediction and legal document generation, the task of judgment review differs fundamentally in its objectives and paradigm: it centers on detecting, classifying, and correcting errors after a judgment is issued, constituting anomaly detection rather than prediction or generation. To address this research gap, we introduce a novel task APPELLATE REVIEW, aiming to assess models' diagnostic reasoning and reliability in legal practice. We also construct a novel dataset benchmark AR-BENCH, which comprises 8,700 finely annotated decisions and 34,617 supplementary corpora. By evaluating 14 large language models, we reveal critical limitations in existing models' ability to identify legal application errors, providing empirical evidence for future improvements.

法律AI错误检测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。