在丢包信道中,用语义重要性自适应编码查询,提升文档检索准确率。
Context-Aware Search and Retrieval Over Erasure Channels
- 根据词语上下文重要性动态分配冗余,编码查询向量
- 理论推导出检索错误率公式,与仿真结果一致
- 适合通信环境差时的智能搜索系统设计
本文提出一种基于语义通信思想的远程文档检索模型,针对符号擦除信道进行信息论分析。模型通过词频权重生成查询特征向量,并采用重复码进行编码,其编码速率依据词语的上下文重要性自适应调整。在解码端,基于恢复后的查询与两文档的上下文相似度进行选择。通过联合高斯近似真实与重构的相似度得分,推导出检索错误概率的显式表达式,即选错文档的概率。合成数据与真实数据(Google NQ)的数值仿真验证了分析有效性。结果表明,对关键特征分配更多冗余可显著降低错误率,证明语义感知编码在易错通信场景中的有效性。
原文摘要 · Abstract (English)
This paper introduces and analyzes a search and retrieval model that adopts key semantic communication principles from retrieval-augmented generation. We specifically present an information-theoretic analysis of a remote document retrieval system operating over a symbol erasure channel. The proposed model encodes the feature vector of a query, derived from term-frequency weights of a language corpus by using a repetition code with an adaptive rate dependent on the contextual importance of the terms. At the decoder, we select between two documents based on the contextual closeness of the recovered query. By leveraging a jointly Gaussian approximation for both the true and reconstructed similarity scores, we derive an explicit expression for the retrieval error probability, i.e., the probability under which the less similar document is selected. Numerical simulations on synthetic and real-world data (Google NQ) confirm the validity of the analysis. They further demonstrate that assigning greater redundancy to critical features effectively reduces the error rate, highlighting the effectiveness of semantic-aware feature encoding in error-prone communication settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。