构建首个支持多层级法律检索的精准引用数据集
LegalPincite: Multi-level Legal Information Retrieval Dataset

- 从欧盟法院判决中提取全段落文本,去除引用信息生成查询
- 包含案件级与段落级真实引用标注,支持跨层级检索任务
- 解决数据泄露问题,适合训练评估真实法律检索系统
法律信息检索(IR)的常见任务是从判例库中找到相关法律来源。尽管法律实务常需定位到具体判例段落的精确引用(pincites),但现有公开法律IR数据集大多缺乏段落级引用标注。而少数含此类信息的数据集存在查询文本中的数据泄露问题,并排除既不被引用也不引用他人的段落,造成过于理想化的检索设定,可能导致性能虚高。为解决上述局限,我们构建了一个大规模法律IR数据集,基于欧盟法院(CJEU)判决。该数据集包含:(i) 去除引用信息的掩码案件/段落查询;(ii) 包含所有段落的完整语料库;(iii) 经部分专家人工验证的案件级与段落级真实引用标注。本数据集支持多层级检索方法的研发与严格评估,涵盖案件-案件、段落-案件、段落-段落三种检索任务。数据集链接:https://huggingface.co/datasets/theresiavr/legalpincite
原文摘要 · Abstract (English)
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground-truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset: https://huggingface.co/datasets/theresiavr/legalpincite
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。