arXiv:2605.21059cs.CVcs.LG2026-05

用成对模态训练多模态大模型,省去复杂对齐数据。

Multimodal LLMs under Pairwise Modalities

论文配图:Multimodal LLMs under Pairwise Modalities
图 1 · 摘自论文原文
  • 仅用成对模态数据学习共享潜在空间,避免全模态对齐。
  • 在3个模态对上实现强跨模态性能,支持3D点云与触觉模态扩展。
  • 适合想低成本扩展多模态能力的研究者和开发者。

尽管多模态大语言模型(MLLMs)取得了显著成果,其训练通常依赖于精心构建的多模态对齐数据集,需大量人力投入,限制了跨领域的可扩展性。本文探索仅利用多个成对模态作为完整联合模态分布的替代方案。我们首先从理论上分析仅观测成对模态时表征可识别的条件。基于此,提出一种仅使用成对数据的表征学习框架,包含两阶段:潜空间对齐与跨模态重构。第一阶段通过自模态重建和成对对比学习,结合归纳偏置(部分对齐与最小潜表征约束)学习跨模态共享潜空间;第二阶段将新引入模态的编码器与预训练模态解码器融合,实现跨模态迁移与生成。我们在三个模态对上新增3D点云和触觉模态,验证该方法通过学习对齐潜空间,实现了优异的跨模态表现。

原文摘要 · Abstract (English)

Despite the impressive results achieved by multimodal large language models (MLLMs), their training typically relies on jointly curated multimodal data, requiring substantial human effort to construct multi-way aligned datasets and thereby limiting scalability across domains. In this work, we explore training MLLMs by only leveraging multiple paired modalities as a surrogate for the full joint multimodal distribution. Specifically, we first provide a theoretical analysis of the conditions under which the representations are identifiable with only observing pairwise modalities. Building on this analysis, we propose a representation learning framework for aligning latent representations across modalities using only pairwise data. The framework consists of two stages: latent representation alignment and cross-modal recomposition. Specifically, in the first stage, we learn the shared latent space across modalities by both self-modal reconstruction and pair-wise contrastive learning. We also incorporate an inductive bias in the contrastive learning process by partially aligning and minimal latent specification. In stage two, we integrate the encoder of newly introduced modalities with the decoders of the pre-trained modalities to facilitate cross-modal transfer and generation. We evaluate our method by newly adding 3D point clouds and tactile modalities into pre-trained MLLMs with three modality pairs and show that, by learning an aligned latent representation space, our model achieves strong cross-modal performance.

多模态大模型表示学习跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。