用生成式模型把法律案件检索变成法律要素推理,更准且泛化更强。
LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation
- 基于查询案情生成关键法律要素,通过约束解码对齐真实判例
- 在120万案例库上提升6-20%,超越现有最佳方法
- 首个大规模韩语法律检索基准,适合司法科技与AI法学研究者
法律案件检索(LCR)是法律从业者研究与决策的基础任务。然而现有研究存在两大局限:其一,评估数据集规模小(如100-55,000个案例),查询类型狭窄,难以反映真实场景复杂性;其二,依赖嵌入或词法匹配,常导致表征有限且匹配不具法律相关性。为此,本文提出:(1) LEGAR BENCH,首个大规模韩语法律案件检索基准,覆盖411种犯罪类型,含120万候选案例;(2) LegalSearchLM,一种在查询案件上进行法律要素推理并直接生成包含这些要素内容的检索模型,通过约束解码确保生成内容与目标案例一致。实验表明,LegalSearchLM在LEGAR BENCH上相较基线提升6%-20%,达到当前最优性能,并在跨领域案例上表现出强泛化能力,优于仅在域内训练的生成模型15%。
原文摘要 · Abstract (English)
Legal Case Retrieval (LCR), which retrieves relevant cases from a query case, is a fundamental task for legal professionals in research and decision-making. However, existing studies on LCR face two major limitations. First, they are evaluated on relatively small-scale retrieval corpora (e.g., 100-55K cases) and use a narrow range of criminal query types, which cannot sufficiently reflect the complexity of real-world legal retrieval scenarios. Second, their reliance on embedding-based or lexical matching methods often results in limited representations and legally irrelevant matches. To address these issues, we present: (1) LEGAR BENCH, the first large-scale Korean LCR benchmark, covering 411 diverse crime types in queries over 1.2M candidate cases; and (2) LegalSearchLM, a retrieval model that performs legal element reasoning over the query case and directly generates content containing those elements, grounded in the target cases through constrained decoding. Experimental results show that LegalSearchLM outperforms baselines by 6-20% on LEGAR BENCH, achieving state-of-the-art performance. It also demonstrates strong generalization to out-of-domain cases, outperforming naive generative models trained on in-domain data by 15%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。