提出统计方法动态选择嵌入维度,显著压缩大小且保持检索效果。
Statistical Foundations of DIME: Risk Estimation for Practical Index Selection
- 基于统计准则,在推理时动态筛选每条查询的最优维度
- 平均减少50%嵌入大小,检索效果与原方法相当
- 适合需要高效部署高维嵌入的检索系统
高维稠密嵌入已成为现代信息检索的核心,但许多维度存在噪声或冗余。近期提出的DIME(维度重要性估计)能为查询提供依赖查询的得分,以识别嵌入中信息量高的部分。然而,DIME需通过代价高昂的网格搜索预先确定所有查询嵌入的维度数。本文提出一种基于统计学的准则,可在推理时直接为每个查询识别出最优维度集合。实验表明,该方法在不同模型和数据集上实现了与原方法相当的检索效果,同时平均将嵌入大小缩减约50%。
原文摘要 · Abstract (English)
High-dimensional dense embeddings have become central to modern Information Retrieval, but many dimensions are noisy or redundant. Recently proposed DIME (Dimension IMportance Estimation), provides query-dependent scores to identify informative components of embeddings. DIME relies on a costly grid search to select a priori a dimensionality for all the query corpus's embeddings. Our work provides a statistically grounded criterion that directly identifies the optimal set of dimensions for each query at inference time. Experiments confirm achieving parity of effectiveness and reduces embedding size by an average of $\sim50\%$ across different models and datasets at inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。