arXiv:2605.17691cs.CLcs.AI2026-05中稿 · publication at the…被引 3

评测大模型在法律判例多标签分类中的表现,提出更贴合实际的评估方法。

Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

论文配图:Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification
图 1 · 摘自论文原文
  • 构建239条真实法律引用数据集,引入严重性误差指标衡量错误影响。
  • Gemini 2.5 Flash在粗粒度任务上准确率79.1%,GPT-5-mini在细粒度任务上达67.7%。
  • 为法律文本理解提供可复现基准,适合法律AI与NLP研究者参考。

自动化法律判例中负面治疗措施的分类是一项关键但复杂的自然语言处理任务,误分类可能带来重大风险。针对传统准确率评估的不足,本文提出更稳健的评估框架。我们在一个由专家标注的、包含239条真实法律引用的新数据集上对现代大语言模型进行基准测试,并引入一种新的平均严重性误差(Average Severity Error)指标,以更好衡量分类错误的实际影响。实验结果揭示性能差异:在高层次分类任务中,Google的Gemini 2.5 Flash表现最佳,准确率达79.1%;而在更复杂的细粒度分类中,OpenAI的GPT-5-mini表现最优,准确率为67.7%。本工作建立了重要基准,提供了上下文丰富的数据集,并提出了契合该复杂法律推理任务的评估指标。

原文摘要 · Abstract (English)

Automating the classification of negative treatment in legal precedent is a critical yet nuanced NLP task where misclassification carries significant risk. To address the shortcomings of standard accuracy, this paper introduces a more robust evaluation framework. We benchmark modern Large Language Models on a new, expert-annotated dataset of 239 real-world legal citations and propose a novel Average Severity Error metric to better measure the practical impact of classification errors. Our experiments reveal a performance split. Google's Gemini 2.5 Flash achieved the highest accuracy on a high-level classification task (79.1%), while OpenAI's GPT-5-mini was the top performer on the more complex fine-grained schema (67.7%). This work establishes a crucial baseline, provides a new context-rich dataset, and introduces an evaluation metric tailored to the demands of this complex legal reasoning task.

法律AI大模型评测多标签分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。