arXiv:2506.09344cs.AIcs.CL2025-06被引 71

一个能统一处理图文音视频的开源模型,还能生成高质量语音和图像。

Ming-Omni: A Unified Multimodal Model for Perception and Generation

  • 用专用编码器提取多模态特征,通过新路由器实现高效融合
  • 支持语音与图像生成,效果媲美GPT-4o,且可直接用于对话编辑
  • 首个开源全模态统一模型,适合研究多模态生成与交互系统者

我们提出Ming-Omni,一个统一的多模态模型,可处理图像、文本、音频和视频,并在语音与图像生成方面表现优异。Ming-Omni采用专用编码器提取各模态的令牌,由名为Ling的MoE架构及新提出的模态专用路由机制进行处理。该设计使单一模型能在统一框架内高效处理与融合多模态输入,无需额外模型、任务微调或结构重设计即可完成多样化任务。尤为重要的是,Ming-Omni突破传统多模态模型局限,支持音频与图像生成:通过先进音频解码器实现自然语音输出,结合Ming-Lite-Uni实现高质量图像生成,同时支持上下文感知对话、文本转语音及多种图像编辑。实验表明,Ming-Omni在所有模态的统一感知与生成任务中表现强大。值得注意的是,Ming-Omni是我们所知首个在模态支持上匹配GPT-4o的开源模型,已公开全部代码与模型权重,以推动社区研究与发展。

原文摘要 · Abstract (English)

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. Ming-Omni employs dedicated encoders to extract tokens from different modalities, which are then processed by Ling, an MoE architecture equipped with newly proposed modality-specific routers. This design enables a single model to efficiently process and fuse multimodal inputs within a unified framework, thereby facilitating diverse tasks without requiring separate models, task-specific fine-tuning, or structural redesign. Importantly, Ming-Omni extends beyond conventional multimodal models by supporting audio and image generation. This is achieved through the integration of an advanced audio decoder for natural-sounding speech and Ming-Lite-Uni for high-quality image generation, which also allow the model to engage in context-aware chatting, perform text-to-speech conversion, and conduct versatile image editing. Our experimental results showcase Ming-Omni offers a powerful solution for unified perception and generation across all modalities. Notably, our proposed Ming-Omni is the first open-source model we are aware of to match GPT-4o in modality support, and we release all code and model weights to encourage further research and development in the community.

多模态生成模型开源语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。