arXiv:2409.05929cs.LGcs.AI2024-09被引 13

用多门专家网络提升跨模态对齐,解决模态坍缩问题。

M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive Architecture

  • 基于JEPA架构,在隐空间中通过多门专家网络实现跨模态对齐
  • 在多个数据集上达到顶尖性能,且在未见数据上表现稳定
  • 适合追求高效自监督学习的科研与工业应用

当前多模态学习主要在原始标记空间优化,虽易接入预训练语言模型,但易导致模态坍缩。为此,本文采用联合嵌入预测架构(JEPA),通过预测器将输入嵌入映射到输出嵌入空间,并在潜在空间进行跨模态对齐。预测器采用多门混合专家(MMoE)结构,门控函数分离模态特有与共享信息,实现信息论最优。框架同时使用对比损失与正则化损失,通过不同模态任务间的交替梯度下降求解。大量实验表明,M3-JEPA在多种模态与任务上均达先进水平,具备良好的跨数据集与跨域泛化能力,且训练与推理效率高。结果表明,该方法可能成为开放世界自监督学习的新基础。

原文摘要 · Abstract (English)

Current multimodal learning strategies primarily optimize in the original token space. Such a framework is easy to incorporate with the backbone of pretrained language model, but might result in modality collapse. To alleviate such issues, we leverage the Joint-Embedding Predictive Architecture (JEPA) on the multimodal tasks, which converts the input embedding into the output embedding space by a predictor and then conducts the cross-modal alignment on the latent space. We implement this predictor by a Multi-Gate Mixture of Experts (MMoE) and name the framework as M3-JEPA, accordingly. The gating function disentangles the modality-specific and shared information and derives information-theoretic optimality. The framework is implemented with both contrastive and regularization loss, and solved by alternative gradient descent (AGD) between different multimodal tasks. By thoroughly designed experiments, we show that M3-JEPA can obtain state-of-the-art performance on different modalities and tasks, generalize to unseen datasets and domains, and is computationally efficient in both training and inference. Our observation suggests that M3-JEPA might become a new basis to self-supervised learning in the open world.

多模态自监督专家网络对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。