arXiv:2608.21494cs.IRcs.DB2026-08

证明单向量检索模型在某些场景下需指数级规模,而多向量只需多项式规模。

Retrieval Needs Multivectors: An Exponential Separation

  • 构建首个理论明确的查询-文档集,揭示单向量与多向量的表达力差距。
  • 单向量模型在ANDOR基准上零样本表现差,微调后提升有限。
  • 多向量模型性能显著优于单向量,且微调后接近理论预测效果。

近期研究通过理论分析和如LIMIT等挑战性基准揭示了基于嵌入的检索模型的表达局限性。尽管多向量嵌入始终优于单向量嵌入,但两者间的具体表征差距仍不清晰。本文沿袭Jayaram的工作,首次提供了明确的查询与文档集合及其相关性矩阵,其中单向量嵌入需指数规模才能将所有相关文档排在无关文档之前,而多项式规模的多向量嵌入即可实现。该结果建立了单向量与多向量嵌入在排序任务中的指数级表达力差距,区别于Jayaram工作中对数值分数近似的设定。受此理论构造启发,我们提出新检索基准ANDOR,自然生成这些困难样本。结果显示,当前最先进的单向量嵌入模型在ANDOR零样本设置下表现不佳,微调后仅略有改善,凸显其固有难度;相反,多向量模型始终表现更优,且微调后大幅提升,与理论预测高度一致。

原文摘要 · Abstract (English)

Recent works have highlighted the expressive limitations of embedding based retrieval models through both theoretical analyses and challenging benchmarks such as LIMIT. While multi-vector embeddings consistently outperform single-vector embeddings, the precise representational gap between them remains poorly understood. In this work, following Jayaram's work, we provide the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice. Our result establishes an exponential separation between the expressive power of single-vector and multi-vector embeddings for the task of ranking of documents as opposed to approximating numerical scores as in the work of Jayaram. Motivated by our theoretical construction, we introduce ANDOR, a new retrieval benchmark that naturally instantiates these hard examples. We show that state-of-the-art single-vector embedding models perform poorly on ANDOR in the zero-shot setting and exhibit only marginal improvements after fine-tuning, highlighting the inherent difficulty of the benchmark compared to prior work. In contrast, multi-vector models consistently outperform their single-vector counterparts and improve substantially with fine-tuning, closely aligning with our theoretical predictions.

检索嵌入理论分析多向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。