首个统一图文音视频的扩散模型,实现多模态理解与生成。
Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

- 用统一离散空间的掩码扩散机制建模多模态数据。
- 在19项评测中多项指标超越开源模型,如GSM8K达87.6。
- 适合需要实时多模态交互的系统或智能体应用。
我们提出Dynin-Omni,首个基于掩码扩散的多模态基础模型,统一文本、图像、语音和视频的理解与生成。不同于串行化处理的自回归模型或需外部解码器的组合式模型,Dynin-Omni将多模态建模原生表述为共享离散令牌空间上的掩码扩散,支持双向上下文下的迭代优化。采用分阶段训练策略,结合模型合并与多模态对齐。我们在涵盖语言推理、图像生成与编辑、视频理解、语音识别与合成的19个基准上评估该模型,取得GSM8K 87.6、MME-P 1733.6、VideoMME 61.4、GenEval 0.87、LibriSpeech test-clean WER 2.1的成绩,持续领先现有开源统一模型,并保持与强模态专精系统相当的竞争力。结果表明,掩码扩散是任意模态间建模的潜在统一范式,为实时多模态系统、跨模态检索与生成、具身多模态智能体提供灵活基础。
原文摘要 · Abstract (English)
We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive unified models that serialize heterogeneous modalities, or compositional unified models that require orchestration with external modality-specific decoders, Dynin-Omni natively formulates omnimodal modeling as masked diffusion over a shared discrete token space, enabling iterative refinement under bidirectional context. Dynin-Omni adopts a multi-stage training strategy with model-merging-based modality expansion and omnimodal alignment. We evaluate Dynin-Omni across 19 multimodal benchmarks spanning language reasoning, image generation and editing, video understanding, and speech recognition and synthesis. Dynin-Omni achieves 87.6 on GSM8K, 1733.6 on MME-P, 61.4 on VideoMME, 0.87 on GenEval, and 2.1 WER on LibriSpeech test-clean, consistently outperforming existing open-source unified models while remaining competitive with strong modality-specific expert systems. These results demonstrate the potential of masked diffusion as a unified paradigm for any-to-any modeling, providing a flexible foundation for real-time omnimodal systems, unified cross-modal retrieval and generation, and embodied multimodal agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。