arXiv:2606.04280cs.LGcs.AI2026-06

对比学习需多样采样才能恢复真实特征结构,否则模型会受先验偏见影响。

The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning

论文配图:The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning
图 1 · 摘自论文原文
  • 提出测度论框架,证明正样本采样需满足多样性条件才可恢复特征几何
  • 在受限采样下,非正交映射能获得更低损失,导致解不唯一
  • 新方法修正了信息噪声对比估计,适合采样不足场景,对架构偏见更敏感

对比学习已成为自监督表示学习的主流范式,但其恢复有意义潜在几何的条件仍不明确。本文建立测度论框架,形式化了多样性条件——即正样本采样需满足支撑要求,这是实现等距潜在空间恢复的必要条件。我们证明标准全支撑冯·米塞斯-费舍尔设置满足该条件,此时全局对比损失最小化器可恢复潜在几何(至正交变换)。而受限条件分布可能导致非正交映射达到更低渐近损失。为此,我们提出支持修正的信息噪声对比估计(InfoNCE)变体作为理论修复:虽能实现正交恢复,但无法唯一确定。合成基准实验验证了可识别性预测;在CIFAR-10上的实验与定性预测一致:当采样多样性受限时,网络架构的归纳偏置变得更为重要。结果揭示了采样机制与编码器归纳偏置在对比学习中的交互作用。

原文摘要 · Abstract (English)

Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful latent geometry remain incompletely understood. We develop a measure-theoretic framework formalizing the diversity condition, a support requirement on positive-pair sampling that is necessary for isometric latent recovery. We show that the standard full-support von Mises-Fisher setting implies the satisfaction of the diversity condition and as a consequence global contrastive loss minimizers recover latent geometry up to orthogonal transformation, while restricted conditionals can make non-orthogonal maps attain strictly lower asymptotic contrastive loss. We introduce a support-corrected Information Noise Contrastive Estimation (InfoNCE) variant as a theoretical fix: this correction makes orthogonal latent space recovery achievable but does not uniquely select it. Experiments on synthetic benchmarks validate the identifiability predictions, and CIFAR-10 experiments are consistent with the qualitative prediction that architectural inductive bias becomes more important when sampling diversity is limited. Together, our results clarify how sampling mechanisms and encoder inductive bias interact in contrastive representation learning.

对比学习自监督表征学习归纳偏置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。