arXiv:2605.07825cs.MMcs.CV2026-05被引 2

提出新方法让单模数据训练多模模型更有效

Anisotropic Modality Align

  • 发现模态间差异是定向的几何结构,非均匀分布
  • 通过有限修正源模态表示,生成目标模态替代表示
  • 适合无配对数据时训练多模大模型的研究者

多模态大语言模型训练长期受限于高质量配对数据稀缺。近期研究发现,预训练多模态对比模型的共享表示空间可作为桥梁,使模型能使用单模数据进行多模训练。然而该范式的前提仍不明确:不同模态的表示能否可靠互换?核心障碍在于共享空间中持续存在的模态差距。本文重新审视模态差距的几何本质,发现模态表示已共享兼容的主导语义几何,真正阻碍互换性的是集中在少数主导方向上的各向异性残差结构。基于此,提出各向异性模态差距对齐原则:有效的模态对齐应匹配目标模态分布,同时保留源模态的语义结构。据此提出AnisoAlign框架,利用目标模态内部几何先验,对源模态表示进行有界修正,从而构建目标模态的替代表示。实验验证其在几何诊断和仅文本的多模大模型训练中的有效性。本工作将模态差距从经验现象重构为可纠正的结构化几何问题,为无配对数据训练多模模型提供了新的表示对齐视角。

原文摘要 · Abstract (English)

Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared representation space of pretrained multimodal contrastive models can serve as a bridge, enabling models to perform multimodal training with unimodal data. However, the key premise of this paradigm remains insufficiently understood: can representations from different modalities be reliably interchanged? The core obstacle lies in the persistent Modality Gap in the shared space. In this work, we revisit the geometric nature of the modality gap. We find that modality representations already share compatible dominant semantic geometry. What truly hinders modality interchangeability is not a simple global shift, but an anisotropic residual structure concentrated along a small number of dominant directions. Based on this finding, we further propose the principle of anisotropic modality gap alignment: effective modality alignment should align with the target-modality distribution while preserving the semantic structure of the source modality. Guided by this principle, we propose an anisotropic geometric correction framework, AnisoAlign, for unpaired modality alignment. This framework leverages the internal geometric prior of the target modality and performs bounded correction on source-modality representations, thereby constructing substitute representations in the target modality. Experiments confirm its benefits in both geometric diagnostics and text-only MLLM training. Overall, this work recasts the modality gap from an empirical observation into a correctable, structured geometric phenomenon and provides a new representation alignment perspective for training multimodal models with unimodal data.

多模态表示对齐无配对数据几何修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。