arXiv:2511.00268cs.CLcs.AI2025-11EMNLP被引 5

构建首个统一法律条文与判例检索的中文语料库

IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval

  • 设计统一语料库,同时支持法律条文与判例检索任务
  • 基于LLM重排序方法在两项任务上均取得最优效果
  • 适合法律AI研究者及智能司法系统开发者

为法律实务中识别/检索相关法律条文和既有判例这一常见任务,现有研究长期将两项任务独立处理,导致数据集与模型体系分离。然而,二者存在内在关联:相似案情通常引用相似条文。本文提出IL-PCSR(Indian Legal corpus for Prior Case and Statute Retrieval),是首个为两类检索任务提供统一测试基准的语料库,支持模型利用二者间的依赖关系。我们在该语料库上对多种基线模型进行实验,涵盖词法、语义模型及基于GNN的集成模型。为进一步挖掘任务间关联性,我们设计一种基于大语言模型的重排序方法,在两项任务中均取得最佳性能。

原文摘要 · Abstract (English)

Identifying/retrieving relevant statutes and prior cases/precedents for a given legal situation are common tasks exercised by law practitioners. Researchers to date have addressed the two tasks independently, thus developing completely different datasets and models for each task; however, both retrieval tasks are inherently related, e.g., similar cases tend to cite similar statutes (due to similar factual situation). In this paper, we address this gap. We propose IL-PCR (Indian Legal corpus for Prior Case and Statute Retrieval), which is a unique corpus that provides a common testbed for developing models for both the tasks (Statute Retrieval and Precedent Retrieval) that can exploit the dependence between the two. We experiment extensively with several baseline models on the tasks, including lexical models, semantic models and ensemble based on GNNs. Further, to exploit the dependence between the two tasks, we develop an LLM-based re-ranking approach that gives the best performance.

法律AI信息检索语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。