arXiv:2510.13253cs.CVcs.AI2025-10ICCV被引 6

用统一架构同时生成高分辨率图像和长文本,提升多模态模型效率

End-to-End Multi-Modal Diffusion Mamba

  • 基于Mamba的扩散模型分步生成并优化多模态信息
  • 在图像生成与文本理解任务中超越现有端到端模型
  • 适合需要高效多模态生成的应用场景

当前端到端多模态模型使用独立编码器与解码器处理输入输出,限制了模态间的联合表征学习。为此,我们提出新型架构MDM(Multi-modal Diffusion Mamba),采用基于Mamba的多步选择扩散模型,通过统一的变分自编码器实现编码与解码,逐步生成并优化特定模态信息。该方法在处理高维数据时表现优异,尤其在同时生成高分辨率图像与长文本序列方面。在图像生成、图像描述、视觉问答、文本理解与推理等任务上的评估表明,MDM显著优于现有端到端模型(如MonoFormer、LlamaGen、Chameleon等),并可与GPT-4V、Gemini Pro、Mistral等顶尖模型比肩。结果验证了MDM在统一多模态流程的同时保持计算高效性,为端到端多模态架构提供了新方向。

原文摘要 · Abstract (English)

Current end-to-end multi-modal models utilize different encoders and decoders to process input and output information. This separation hinders the joint representation learning of various modalities. To unify multi-modal processing, we propose a novel architecture called MDM (Multi-modal Diffusion Mamba). MDM utilizes a Mamba-based multi-step selection diffusion model to progressively generate and refine modality-specific information through a unified variational autoencoder for both encoding and decoding. This innovative approach allows MDM to achieve superior performance when processing high-dimensional data, particularly in generating high-resolution images and extended text sequences simultaneously. Our evaluations in areas such as image generation, image captioning, visual question answering, text comprehension, and reasoning tasks demonstrate that MDM significantly outperforms existing end-to-end models (MonoFormer, LlamaGen, and Chameleon etc.) and competes effectively with SOTA models like GPT-4V, Gemini Pro, and Mistral. Our results validate MDM's effectiveness in unifying multi-modal processes while maintaining computational efficiency, establishing a new direction for end-to-end multi-modal architectures.

多模态生成扩散模型Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。