arXiv:2605.17524cs.LGcs.DB2026-05

揭示对比嵌入二值化成功背后的协方差与坐标异质性机制

Covariance Structure and Coordinate Heterogeneity Govern Binary Quantization of Contrastive Embeddings

  • 基于高斯模型分析协方差结构对排序保真度的影响
  • 发现非对角相关项贡献30%-50%信号,坐标异质性决定比特增益
  • 解释随机旋转与轴保留策略的对立原理,提供首个设计指南

二值量化(BQ)将高维嵌入压缩为每坐标1或2比特,实现极快最近邻搜索。然而存在一个显著谜题:BQ在对比嵌入上表现良好,但在其他嵌入上失败——两种主流系统采用截然相反策略(随机旋转 vs. 保持坐标轴),却缺乏统一理论解释其适用条件。本文通过将InfoNCE训练表示的高斯结构与BQ质量的统计框架相连接,揭示协方差矩阵的双重作用:第一,完整协方差结构(不仅对角线)决定排序保真度的绝对水平,非对角相关项贡献30%-50%信号;第二,坐标异质性(各坐标方差的不均匀性)决定关键设计选择:每增加一个比特的收益,以及随机旋转是否有益。我们推导了高斯模型下的排序保真度近似表达式,证明幅度比特携带的信息量与异质性成正比,并表明随机旋转会破坏一种范式依赖的信号,同时创造另一种范式所需的各向同性。现象级缩放定律可预测跨模型与维度的保真度。18个数据集、9类嵌入家族的实验验证了核心预测,首次提供了二值量化系统的原理性设计指导。

原文摘要 · Abstract (English)

Binary quantization (BQ) compresses high-dimensional embeddings into one or two bits per coordinate, enabling nearest neighbor search at extreme speed. Yet a striking puzzle persists: BQ achieves competitive recall on contrastive embeddings but fails on others -- and two leading systems adopt diametrically opposite strategies (random rotation vs. preserving coordinate axes) without a common theory explaining when each is appropriate. We address this puzzle by connecting the Gaussian structure recently established for InfoNCE-trained representations to a statistical framework for BQ quality. Our analysis reveals two distinct roles of the covariance matrix. First, the full covariance structure -- not merely its diagonal -- determines the absolute level of ranking fidelity, with off-diagonal correlations contributing 30--50% of the signal. Second, coordinate heterogeneity (the non-uniformity of per-coordinate variances) governs key design choices: how much each additional bit contributes, and whether random rotation helps or hurts. We derive approximate expressions for ranking fidelity under a Gaussian model, show that the magnitude bit carries information proportional to heterogeneity, and show that random rotation destroys precisely the signal that one paradigm exploits while creating the isotropy that the other requires. A phenomenological scaling law predicts fidelity across models and dimensions. Experiments on 18 datasets spanning 9 embedding families support the main predictions and provide, to our knowledge, the first principled design guide for binary quantization systems.

二值量化对比学习协方差分析嵌入压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。