用算法自动生成法律案件重要性标签,提升法院案件优先级判断效率。
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
- 基于引用频率与时效性构建双层标签体系,自动化生成大规模标注数据。
- 小模型经大规模训练后表现优于大模型,证明数据量对专业领域关键。
- 适合法律科技、司法智能化研究者,为案件优先级系统提供新基准。
全球许多法院系统面临案件积压问题,亟需类似急诊分诊的高效优先级系统以优化资源分配。本文提出「Criticality Prediction」数据集,用于评估案件优先级预测能力。该数据集采用两层标注:(1) 二分类的LD-Label,识别是否为判例性判决(Leading Decisions);(2) 更精细的Citation-Label,按引用频次与时效性对案件排序,支持更细致评估。不同于依赖人工标注的现有方法,本研究通过算法自动推导标签,实现远超传统规模的数据集构建。我们评估了多种多语言模型,包括微调的小型模型和零样本设置下的大型语言模型。结果表明,微调模型在所有任务上均显著优于大型模型,归因于其大规模训练数据。这表明,在高度专业化任务中,大规模训练集仍具关键价值。
原文摘要 · Abstract (English)
Many court systems are overwhelmed all over the world, leading to huge backlogs of pending cases. Effective triage systems, like those in emergency rooms, could ensure proper prioritization of open cases, optimizing time and resource allocation in the court system. In this work, we introduce the Criticality Prediction dataset, a novel resource for evaluating case prioritization. Our dataset features a two-tier labeling system: (1) the binary LD-Label, identifying cases published as Leading Decisions (LD), and (2) the more granular Citation-Label, ranking cases by their citation frequency and recency, allowing for a more nuanced evaluation. Unlike existing approaches that rely on resource-intensive manual annotations, we algorithmically derive labels leading to a much larger dataset than otherwise possible. We evaluate several multilingual models, including both smaller fine-tuned models and large language models in a zero-shot setting. Our results show that the fine-tuned models consistently outperform their larger counterparts, thanks to our large training set. Our results highlight that for highly domain-specific tasks like ours, large training sets are still valuable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。