用JAM让独立训练的视觉语言模型对齐,提升细粒度语义理解。
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
- 通过联合训练模态自编码器,实现跨模态对齐。
- 提出多模态扩散损失,优于传统对比方法。
- 适用于将单模态大模型转化为专业多模态模型。
独立训练的视觉与语言模型处于各自分离的表征空间中,受模态、目标和架构影响。柏拉图表征假说(PRH)认为它们可能仍收敛于对现实的共享统计模型。这引发关键问题:能否超越事后检测,主动优化这种对齐?我们聚焦细粒度上下文差异——多个描述具有全局语义相似性但组合细节不同。为此提出联合自编码调制器(JAM),通过联合训练模态特定自编码器,并设置协同重构与跨模态对齐目标,实现冻结单模态模型的对齐。系统评估了三个设计维度:(i) 对齐目标,引入多模态扩散损失,表现优于经典对比方法;(ii) 对齐最有效的层深度;(iii) 基础模型规模对表征收敛的影响。结果表明,即使在独立训练的表示间,JAM也能可靠诱导对齐,既提供共享语义结构的理论洞察,也为将通用单模态基础模型转化为专用多模态模型提供实践指导。
原文摘要 · Abstract (English)
Independently trained vision and language models inhabit disjoint representational spaces, shaped by their respective modalities, objectives, and architectures. The Platonic Representation Hypothesis (PRH) suggests these models may nonetheless converge toward a shared statistical model of reality. This raises a fundamental question: can we move beyond post-hoc detection of such alignment and explicitly optimize for it? We argue this challenge is most critical in fine-grained contextual distinctions-where multiple descriptions share global semantics but differ in subtle compositional details. We address this with the Joint Autoencoder Modulator (JAM), which aligns frozen unimodal models by jointly training modality-specific autoencoders with coordinated reconstruction and cross-modal alignment objectives. We systematically evaluate JAM across three design axes: (i) alignment objectives, introducing our multimodal Spread Loss that outperforms classic contrastive methods; (ii) the layer depth at which alignment is most effective; and (iii) the role of foundation model scale in representational convergence. Our findings show that JAM reliably induces alignment even across independently trained representations, offering both theoretical insight into the structure of shared semantics and practical guidance for transforming generalist unimodal foundations into specialist multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。