研究检索截断下的可行性保持,揭示系统何时能正确返回答案。
Feasibility Preservation under Monotone Retrieval Truncation
- 将检索建模为截断下的可行性问题,分析答案可得的结构条件。
- 证明单调截断下单个查询的可行性可保证有限深度下仍可发现证据。
- 指出非单调截断或无限查询类会导致失败,适合关注检索可靠性研究者。
基于检索的系统通过仅暴露可用证据的截断子集来近似访问语料库。即使相关证据存在于语料中,截断也可能导致兼容证据无法共现,从而引发无法被基于相关性的评估捕获的失败。本文从结构角度研究检索,将查询回答建模为在截断下的可行性问题。我们形式化检索为候选证据集序列,并刻画了极限可行性蕴含有限深度可行性的条件。我们证明,单调截断足以保证单个查询的有限可证性。对于特定查询类别,我们识别出见证证书的有限生成是获得统一检索深度上界所需的额外条件,且该条件必要。我们进一步给出精确反例,展示非单调截断、非有限生成查询类及纯槽位覆盖下的失败情况。这些结果将可行性保持确立为独立于相关性评分或优化的检索正确性标准,并阐明了基于截断检索固有的结构性限制。
原文摘要 · Abstract (English)
Retrieval-based systems approximate access to a corpus by exposing only a truncated subset of available evidence. Even when relevant information exists in the corpus, truncation can prevent compatible evidence from co-occurring, leading to failures that are not captured by relevance-based evaluation. This paper studies retrieval from a structural perspective, modeling query answering as a feasibility problem under truncation. We formalize retrieval as a sequence of candidate evidence sets and characterize conditions under which feasibility in the limit implies feasibility at finite retrieval depth. We show that monotone truncation suffices to guarantee finite witnessability for individual queries. For classes of queries, we identify finite generation of witness certificates as the additional condition required to obtain a uniform retrieval bound, and we show that this condition is necessary. We further exhibit sharp counterexamples demonstrating failure under non-monotone truncation, non-finitely-generated query classes, and purely slotwise coverage. Together, these results isolate feasibility preservation as a correctness criterion for retrieval independent of relevance scoring or optimization, and clarify structural limitations inherent to truncation-based retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。