现有检索模型常找到相关文档却漏掉关键证据,本文揭示这一覆盖缺陷。
Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval

- 设计可控实验评估多条件文档检索的证据覆盖能力
- 最强模型仅35.8%实现全条件首胜,远低于81.1%的黄金文档召回率
- 强调证据完整覆盖是核心瓶颈,适合信息检索与大模型验证研究者
多部分查询的长文档检索不仅要求找到相关文档,更需确保文档包含所有请求证据。本文针对联合条件检索(conjunctive retrieval)中的证据覆盖缺口展开研究,使用n-Clue作为控制测量工具:在2,021份文档上对1,000个查询进行测试,将完全匹配的黄金文档与仅满足部分条件的自然文档配对,并要求完整首次成功需在前10位中先于所有子集结果出现。70种配置下,条件分解使两种密集型骨干模型提升6.8–7.3点,词法-视觉融合提升8.7点,但四种通用重排序器均降低Gold-NDCG;该趋势在四源压力集上重复验证。将一个密集家族从0.6B扩展至8B,完整首次成功率无变化(0.0点)。最强混合模型可为81.1%查询找到黄金文档,但仅35.8%实现完整首次成功,且该差距在条件数、目标长度、候选密度、查询呈现方式及四源压力集上持续存在。页面感知视觉系统仅在5.1–5.3%查询中完整呈现各条件支持。结果表明,证据覆盖而非黄金发现才是核心瓶颈。
原文摘要 · Abstract (English)
Finding a long document relevant to a multi-part request is not the same as establishing that it contains every requested piece of evidence. We study this gap for conjunctive document retrieval, where two or three explicit conditions must be supported on different pages of one document. We use n-Clue as a controlled measurement instrument: 1{,}000 queries over 2{,}021 documents pair all-condition golds with naturally occurring documents that satisfy only a subset, and a complete-first success requires a top-10 gold to precede every released subset qrel. Across 70 configurations, condition-wise decomposition improves two dense backbones by 6.8--7.3 points and lexical--visual fusion adds 8.7, while four generic rerankers all reduce Gold-NDCG; these directions replicate on a four-source stress set. Scaling one dense family from 0.6B to 8B changes complete-first success by 0.0 points. The strongest displayed hybrid illustrates the resulting gap: it finds a gold for 81.1\% of queries but succeeds complete-first on only 35.8\%, and the gap persists across condition count, target length, candidate density, query rendering, and the four-source stress set. Finally, page-aware visual systems surface stored support for every condition on only 5.1--5.3\% of queries. These results identify condition coverage, rather than gold discovery alone, as the central bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。