arXiv:2604.18424cs.IRcs.IT2026-04

为部分关键词丢失的检索场景设计了重要性感知冗余机制

Context-Aware Search and Retrieval Under Token Erasure

  • 根据词频与语义重要性动态分配冗余信息
  • 相似度差距向量收敛至多元高斯分布,可计算误差上界
  • 适用于关键词易丢失的RAG系统,提升检索鲁棒性

本文研究了在查询表示部分丢失({token} erasures)条件下远程文档检索的性能。采用基于词频的特征表示查询,依据特征重要性分配语义自适应冗余,通过加权TF-IDF进行检索。我们从信息论角度分析,发现相似度差距向量收敛于多元高斯分布,从而获得误差概率的显式近似和可计算上界。数值结果验证了理论分析,另在真实数据上的嵌入式检索评估表明,该重要性感知冗余原则同样适用于现代检索流程。总体结果表明,对语义关键特征施加更高冗余能显著提升检索可靠性。

原文摘要 · Abstract (English)

This paper introduces and analyzes a search and retrieval model for RAG-like systems under {token} erasures. We provide an information-theoretic analysis of remote document retrieval when query representations are only partially preserved. The query is represented using term-frequency-based features, and semantically adaptive redundancy is assigned according to feature importance. Retrieval is performed using TF-IDF-weighted similarity. We characterize the retrieval error probability by showing that the vector of similarity margins converges to a multivariate Gaussian distribution, yielding an explicit approximation and computable upper bounds. Numerical results support the analysis, while a separate data-driven evaluation using embedding-based retrieval on real-world data shows that the same importance-aware redundancy principles extend to modern retrieval pipelines. Overall, the results show that assigning higher redundancy to semantically important query features improves retrieval reliability.

检索增强鲁棒检索信息论冗余设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。