研究生成式检索的泛化局限,揭示约束解码与束搜索的理论瓶颈。
Constrained Auto-Regressive Decoding Constrains Generative Retrieval
- 通过约束解码生成文档标识,实现端到端检索。
- 发现误差下界由真实与预测分布的KL散度决定。
- 指出束搜索中边际分布使用不够理想,适合理论研究者。
生成式检索旨在用单一大规模神经网络替代传统搜索索引结构,有望提升效率并无缝集成生成式大语言模型。作为端到端范式,生成式检索采用可学习的可微分索引,通过特定语料的约束解码直接生成文档标识。其在分布外语料上的泛化能力备受关注。本文从约束和束搜索两个关键角度审视自回归生成的内在局限。首先在贝叶斯最优设置下,假设模型精确捕捉所有可能文档的相关性分布;随后对具体语料仅添加语料特定约束进行应用。主要发现为:(i) 约束的影响方面,推导出误差的下界,以真实与模型预测的逐步边际分布之间的KL散度表示;(ii) 束搜索算法在生成过程中使用边际分布并非理想方法。本文旨在深化对自回归解码检索范式泛化能力的理论理解,揭示其局限性,为未来更鲁棒、更通用的生成式检索发展奠定基础。
原文摘要 · Abstract (English)
Generative retrieval seeks to replace traditional search index data structures with a single large-scale neural network, offering the potential for improved efficiency and seamless integration with generative large language models. As an end-to-end paradigm, generative retrieval adopts a learned differentiable search index to conduct retrieval by directly generating document identifiers through corpus-specific constrained decoding. The generalization capabilities of generative retrieval on out-of-distribution corpora have gathered significant attention. In this paper, we examine the inherent limitations of constrained auto-regressive generation from two essential perspectives: constraints and beam search. We begin with the Bayes-optimal setting where the generative retrieval model exactly captures the underlying relevance distribution of all possible documents. Then we apply the model to specific corpora by simply adding corpus-specific constraints. Our main findings are two-fold: (i) For the effect of constraints, we derive a lower bound of the error, in terms of the KL divergence between the ground-truth and the model-predicted step-wise marginal distributions. (ii) For the beam search algorithm used during generation, we reveal that the usage of marginal distributions may not be an ideal approach. This paper aims to improve our theoretical understanding of the generalization capabilities of the auto-regressive decoding retrieval paradigm, laying a foundation for its limitations and inspiring future advancements toward more robust and generalizable generative retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。