用核方法实现对比学习的高效闭式解,训练速度显著提升
UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

- 引入核函数构建相似度权重矩阵,实现全局闭式更新
- 在多种任务上保持强性能,训练效率大幅提升
- 适用于线性与非线性编码器及一对一、多对多对齐
对比学习驱动当前主流多模态模型,但训练过程仍依赖长时间随机优化。我们提出统一框架 UniCon(基于核的高效对比对齐),覆盖线性与非线性编码器,以及一对一与多对多对齐。核心是引入对比相似度权重矩阵 $S(γ)$,实现可证明的闭式全局解,能精确替代小批量反向传播。通过再生核希尔伯特空间(RKHS)视角,揭示了对比对齐与谱方法的内在联系。实验在合成数据、单模态、多模态及零样本任务上验证理论,结果表明 UniCon 在保持泛化能力与优异性能的同时,实现显著的效率提升。
原文摘要 · Abstract (English)
Contrastive objectives power state-of-the-art multimodal models, but their training remains slow, relying on long stochastic optimization. We propose a Unified Framework for Efficient Contrastive Alignment via Kernels (UniCon), which spans linear and nonlinear encoders as well as one-to-one and many-to-many alignments. At its core, UniCon introduces the contrastive similarity weight matrix $S(γ)$, which enables closed-form global solutions that provably replace minibatch back-propagation with exact updates. Through the lens of reproducing kernel Hilbert spaces (RKHS), UniCon provides a kernelized perspective that unifies contrastive alignment and reveals its connection to spectral methods. To validate the theory, we conduct experiments on synthetic, unimodal, multimodal, and zero-shot tasks, demonstrating that UniCon achieves substantial efficiency gains while preserving generality and strong empirical performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。