arXiv:2606.21676cs.IR2026-06

构建法律条文跨引用检索数据集,验证查询是否真需上下文。

CRAwLeR -- Cross-Reference Aware Legal Retrieval

论文配图:CRAwLeR -- Cross-Reference Aware Legal Retrieval
图 1 · 摘自论文原文
  • 基于法律条文跨引用设计上下文依赖的检索任务
  • 丹麦和波兰数据集上最佳召回率55%~59%,仍存明显差距
  • 首次关注检索任务的构念效度,适合法律AI研究者

现有上下文感知段落检索基准多依赖改造任务,难以证明查询真正需要上下文,导致评分解释困难。本文聚焦法律条文中的跨引用现象,提出CRAwLeR,一个针对法律文档中段落检索的上下文依赖问题的可操作化范式。该流水线检测法律跨引用,识别查询候选,将目标段落与其相关上下文关联,利用大模型生成需上下文的查询,并通过对抗性非上下文基线和保证提示进行过滤。我们发布了CRAwLeR-DK(丹麦)和CRAwLeR-PL(波兰)两个数据集,以及一种类Anthropic的强上下文基线。人工分析显示约80%的随机抽样查询确实针对标注目标段落且依赖上下文,失败模式系统且可命名。这些基准难度高但未被解决:在CRAwLeR-DK和CRAwLeR-PL上最佳Recall@10分别为55%和59%。消融与失败分析表明剩余差距源于上下文生成的大模型,而非检索器本身。即使目标段落在前10名,标注的上下文段落也常排名更高。本工作是首个系统考虑构念效度的上下文感知段落检索数据集。

原文摘要 · Abstract (English)

Existing benchmarks for context-aware chunk retrieval rely heavily on repurposed task items and rarely demonstrate that their queries genuinely require context, making score interpretation difficult. We focus on a specific kind of context dependence, legal cross-references, and introduce CRAwLeR, an operationalization of a narrow, well-defined phenomenon: cross-reference-aware context utilization for chunk retrieval in legal documents. Our pipeline detects legal cross-references, identifies query candidates, links target chunks to their relevant context, generates context-demanding queries with an LLM, and filters them through both an adversarial non-contextual baseline and an assurance prompt. We release CRAwLeR-DK and CRAwLeR-PL, Danish and Polish datasets built with this pipeline, alongside a strong Anthropic-style contextualization baseline. Manual analysis finds that approximately 80% of randomly sampled queries genuinely target the labelled target chunk and require context, with failures following systematic and named patterns. The benchmarks are hard but not solved: best Recall@10 reaches 55% on CRAwLeR-DK and 59% on CRAwLeR-PL. Ablation and failure analysis attribute the remaining gap to the contextualising LLM, not the retriever. Even when the target is retrieved in the top ten, labelled context chunks routinely outrank it. We are the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect our results in the light of such a narrow, well-defined phenomenon.

法律AI检索上下文依赖数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。