提出原生多模态建模路线图,系统梳理融合架构与应用范式。
Toward Native Multimodal Modeling: A Roadmap

- 按输入输出对称性划分三类原生模型:多模态→文本、多模态→目标生成、多模态→多模态
- 明确区分中融合、早融合与非原生范式,定义架构原生性标准
- 覆盖从数据、训练到部署的全链路工业级实践,适合构建统一多模态系统的研究者
多模态建模是实现跨模态推理向世界建模演进的关键一步。早期方法主要依赖晚融合,即通过拼接编码器与冻结语言主干网络来构建输出头;近期研究转向原生多模态建模(NMM),强调模态间内在融合以提升性能。尽管潜力巨大,当前原生架构的设计空间仍不清晰。本文提出正式路线图,明确定义架构原生性,区分中融合、早融合与非原生范式。基于输入-输出对偶性,将现有原生模型归为三类:(i) 多模态→文本,用于跨模态理解;(ii) 多模态→目标,面向图像、音频、视频等场景化生成;(iii) 多模态→多模态,实现对称统一建模。本工作系统剖析通往终极原生多模态框架的路径,涵盖工业级视角下的架构协同、海量数据治理、端到端训练方案、推理与部署策略及综合评估体系,推动理解与生成在统一Transformer范式中无缝共存。
原文摘要 · Abstract (English)
Multimodal modeling represents a vital step from modality-agnostic reasoning toward world modeling. While early approaches predominantly rely on late-fusion that assembles encoders and frozen language backbones with output heads, recent efforts have shifted the paradigm toward native multimodal modeling (NMM) with the intrinsic integration of modalities for superior multimodal performance. Despite its potential, the design space of native architectures remains insufficiently defined. In this paper, we present the community with a formalized roadmap for this transition. Specifically, we formally define the architectural nativity, distinguishing mid-fusion and early-fusion from non-native paradigms. We further organize the existing native models through the lens of input-output duality into three categories: (i) Multi-to-Text for cross-modal comprehension with text-only output; (ii) Multi-to-Target for scenario-oriented generation, e.g., image, audio and video generation, and (iii) Multi-to-Multi for unified modeling with symmetric input-output. We deliver a comprehensive and industrial-grade investigation into the transition toward the definitive NMM framework, where understanding and generation seamlessly coexist within a unified transformer paradigm. We systematically unpack the end-to-end pipeline from industrial perspectives from architectural coordination, massive data curation, to full-stack training recipes, inference & deployment, and the comprehensive evaluation for truly native modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。