arXiv:2608.30929cs.CL2026-08

用语言模型标注法律条文,提升波兰法律检索精度

Annotated Surrogate Retrieval for Polish Statutory Law

  • 在法律条文上附加语言模型注释,构建可高效检索的代理索引
  • ASCR-H 在前五名召回率上显著优于传统方法,达72.3%
  • 轻量级方案DTF延迟仅为1/9,成本更低,适合实际部署

我们提出一组基于文档代理的波兰法律条文检索方法,通过在索引时为法律条文附加语言模型注释实现。三种设计分别位于成本-质量权衡的不同位置:ASCR采用代理级联与重排序;ASCR-H将稠密检索结果融合进级联;DTF则完全替换语言模型阶段,使用三个词法与稠密检索器,加权倒数排名融合及确定性重评分先验,生成前不调用任何模型。在包含82,508个条文、覆盖1,133部法律的语料库上,针对2024和2025年波兰律师及法律顾问考试的300个问题(其中264个参考条文存在于语料库中)进行评估。配对麦克尼马尔检验显示,除一个自身消融实验外,ASCR-H在将参考条文排在首位方面显著优于所有其他非人工基准(20次比较中有18次显著,p < 0.005),达到72.3%,远超BM25的61.7%和稠密检索的52.3%。优势集中在头部,十名后不再显著;至二十名时,DTF以86.0%领先于ASCR-H的84.5%,且延迟仅为1/9,成本不足一半。消融分析表明,重排序阶段贡献了27.6个百分点的一阶准确率。进一步发现,该排名优势未延伸至引用准确性,此时DTF达到与人工标注天花板相同的86.0%。此外,词干化、伪相关反馈与查询重写均未带来正向效果。代理注释覆盖27.0%的语料库,但所有基准测试中的参考条文均被覆盖,这一不对称性被披露并讨论。评测数据、每题输出及配对显著性检验均已公开。

原文摘要 · Abstract (English)

We present a family of retrieval methods for Polish statutory law built on document surrogates: language-model annotations attached to statutory articles at index time. Three designs occupy different points on the cost-quality frontier. ASCR is a surrogate cascade with reranking; ASCR-H fuses a dense list into that cascade; and DTF replaces both language-model stages with three lexical and dense retrievers, weighted reciprocal rank fusion, and a deterministic re-scoring prior, using no model call before generation. We evaluate all three against fourteen lexical, dense, fused and ablated baselines plus four controls, on 300 questions from the 2024 and 2025 Polish bar and legal counsel entrance examinations (264 with their reference article in the corpus), over 82,508 articles from 1,133 acts. On paired McNemar tests, ASCR-H places the reference provision at rank one significantly more often than every other non-oracle configuration except one of its own ablations (eighteen of twenty comparisons significant in its favour at p < 0.005), reaching 72.3% against 61.7% for BM25 and 52.3% for dense retrieval. The advantage is concentrated at the head and does not survive depth: it is significant at cutoffs of one and five, disappears by ten, and by twenty DTF leads on point estimate (86.0% versus 84.5%) at one ninth the latency and less than half the cost. Ablation attributes 27.6 points of rank-one accuracy to the reranking stage alone. We further report that the ranking advantage does not extend to citation accuracy, where DTF matches the oracle ceiling, and three negative results on lemmatisation, pseudo-relevance feedback and query rewriting. Surrogate annotation covers 27.0% of the corpus but every reference provision in the benchmark, an asymmetry we disclose and discuss. Benchmark, per-question outputs and paired significance tests are publicly available.

法律检索代理检索语言模型波兰法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。