arXiv:2510.13721cs.CLcs.AI2025-10被引 20

首个支持任意模态互转的统一多模态模型,提升交互与检索效率。

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

  • 用离散流匹配实现跨模态统一建模,打破生成与理解的架构壁垒。
  • 在多轮交互和跨模态检索上超越现有统一模型,响应更高效。
  • 开源代码与模型,支持研究者复现与扩展,适合多模态系统开发者。

下一代多模态基础模型需支持任意模态间的跨模态生成与多轮交互,成为通用人工智能系统的核心。然而,现有模型多受限于自回归架构,难以平衡理解与生成能力。虽有混合与解耦策略尝试统一框架,但冗余设计限制了其在跨模态检索等场景的应用。本文提出 NExT-OMNI,一个开源的全模态基础模型,通过离散流范式实现统一建模。借助度量诱导概率路径与动能最优速度,原生支持任意模态间的理解与生成,提升响应效率,并以简洁统一表征拓展应用场景。在大规模交错文本、图像、视频、音频数据上训练,该模型在多模态生成与理解基准上表现优异,且在多轮交互与跨模态检索任务中优于先前统一模型,凸显其架构优势。为推动研究,我们公开训练细节、数据协议以及代码与模型检查点。

原文摘要 · Abstract (English)

Next-generation multimodal foundation models capable of any-to-any cross-modal generation and multi-turn interaction will serve as core components of artificial general intelligence systems, playing a pivotal role in human-machine interaction. However, most existing multimodal models remain constrained by autoregressive architectures, whose inherent limitations prevent a balanced integration of understanding and generation capabilities. Although hybrid and decoupling strategies have been explored to address these tasks within unified frameworks separately, their redundant, non-integrated designs limit their applicability to broader scenarios, such as cross-modal retrieval. In this work, we introduce NExT-OMNI, an open-source omnimodal foundation model that achieves unified modeling through discrete flow paradigms. By leveraging metric-induced probability paths and kinetic optimal velocities, NExT-OMNI natively supports any-to-any understanding and generation with enhanced response efficiency, while enabling broader application scenarios through concise unified representations rather than task-decoupled designs. Trained on large-scale interleaved text, image, video, and audio data, NExT-OMNI delivers competitive performance on multimodal generation and understanding benchmarks, while outperforming prior unified models in multi-turn multimodal interaction and cross-modal retrieval, highlighting its architectural advantages as a next-generation multimodal foundation model. To advance further research, we release training details, data protocols, and open-source both the code and model checkpoints.

多模态统一建模流匹配开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。