解释了对比学习中嵌入长度为何隐含语义信息。
Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms

- 通过优化动态分析,发现嵌入长度是训练的副产品。
- 嵌入长度与概念具体性等语义属性显著相关。
- 可作免费校准信号,适用于检索等任务。
使用尺度不变损失函数训练的对比嵌入模型通常搭配余弦相似度等距离度量,忽略嵌入向量的模长。然而,实证研究发现,这些被忽略的模长仍与概念具体性、词频及人类不确定性等语义属性相关。本文通过分析优化动态,推导出一个解析公式,证明嵌入长度会作为训练过程的副产物自然编码此类信息。我们还展示了该特性可为特定模型和检索任务提供无需额外标注的‘免费’校准信号,为此前的启发式观察提供了理论依据。
原文摘要 · Abstract (English)
Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes. However, surprisingly, empirical studies reveal that despite this, these "discarded" norms seem to correlate with semantic properties such as concept specificity, token frequency, and human uncertainty. In this work, we provide a formal theoretical framework explaining this phenomenon. By analyzing the optimization dynamics, we derive an analytic formula demonstrating that embedding length naturally encodes this information as a byproduct of the training process. We also show how this gives rise to signals that can serve as "free" calibration tools in specific models and retrieval tasks, providing a grounded explanation for a previously heuristic observation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。