揭示对比学习中温度退火的理论机制,指导高效调参
Asymptotic and Finite-Time Guarantees for Langevin-Based Temperature Annealing in InfoNCE
- 用李群流形上的朗之万动力学建模嵌入演化
- 慢速退火可收敛到全局最优表示,快退火易陷局部最优
- 为对比学习的温度调度提供理论依据,适合算法设计者
对比学习中的InfoNCE损失对温度参数极为敏感,但固定与退火温度策略的动态特性尚不明确。本文通过在紧致黎曼流形上建模嵌入演化,基于温和的光滑性与能垒假设,证明经典模拟退火的保证可推广至该场景:缓慢的对数逆温度退火可确保以概率收敛至全局最优表示集,而过快的退火可能导致陷入次优极小值。结果建立了对比学习与模拟退火之间的联系,为理解与优化温度调度提供了理论基础。
原文摘要 · Abstract (English)
The InfoNCE loss in contrastive learning depends critically on a temperature parameter, yet its dynamics under fixed versus annealed schedules remain poorly understood. We provide a theoretical analysis by modeling embedding evolution under Langevin dynamics on a compact Riemannian manifold. Under mild smoothness and energy-barrier assumptions, we show that classical simulated annealing guarantees extend to this setting: slow logarithmic inverse-temperature schedules ensure convergence in probability to a set of globally optimal representations, while faster schedules risk becoming trapped in suboptimal minima. Our results establish a link between contrastive learning and simulated annealing, providing a principled basis for understanding and tuning temperature schedules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。