arXiv:2410.06581cs.IR2024-10EMNLP被引 18

构建超大规模合成法律案例检索数据集,提升模型效果

Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs

  • 自动生成高质量查询-候选对,解决数据稀缺问题
  • 新数据集LEAD规模达现有数据数百倍,支持深度学习训练
  • 适用于民事等多类型案件,可直接用于司法辅助系统

法律案例检索(LCR)旨在为给定案情描述提供相似案例参考,对保障判决一致性、提升司法公平性和法官工作效率至关重要。然而现有研究面临两大挑战:主流方法依赖长篇查询进行案例间匹配,与真实场景不符;且数据规模有限,现有数据集仅含数百个查询,难以满足现代神经模型的训练需求。为此,本文提出一种自动化合成查询-候选对的方法,构建迄今最大的LCR数据集LEAD,其规模为现有数据集的数百倍。该方法为LCR模型提供了充足训练信号。实验表明,基于该数据训练的模型在两个主流基准上达到领先性能。此外,该方法可拓展至民事案件,表现良好。数据与代码见https://github.com/thunlp/LEAD。

原文摘要 · Abstract (English)

Legal case retrieval (LCR) aims to provide similar cases as references for a given fact description. This task is crucial for promoting consistent judgments in similar cases, effectively enhancing judicial fairness and improving work efficiency for judges. However, existing works face two main challenges for real-world applications: existing works mainly focus on case-to-case retrieval using lengthy queries, which does not match real-world scenarios; and the limited data scale, with current datasets containing only hundreds of queries, is insufficient to satisfy the training requirements of existing data-hungry neural models. To address these issues, we introduce an automated method to construct synthetic query-candidate pairs and build the largest LCR dataset to date, LEAD, which is hundreds of times larger than existing datasets. This data construction method can provide ample training signals for LCR models. Experimental results demonstrate that model training with our constructed data can achieve state-of-the-art results on two widely-used LCR benchmarks. Besides, the construction method can also be applied to civil cases and achieve promising results. The data and codes can be found in https://github.com/thunlp/LEAD.

法律AI案例检索合成数据司法辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。