arXiv:2605.12500cs.CV2026-05被引 19

统一视觉语言理解与生成,构建原生多模态智能新范式。

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

论文配图:SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
图 1 · 摘自论文原文
  • 基于NEO-unify架构,将理解和生成视为同一过程的协同视角。
  • 8B和30B-A3B版本在理解与生成任务上均达到顶尖水平。
  • 适合研究多模态统一架构、生成与推理融合的开发者与学者。

当前大型视觉语言模型(VLMs)仍受制于理解与生成分离的根本矛盾,导致架构碎片化、流水线级联及表征空间错位。我们提出SenseNova-U1,一种基于NEO-unify的原生统一多模态范式,使理解与生成成为单一底层过程的协同视图。推出两个原生统一变体:SenseNova-U1-8B-MoT(密集型8B)与SenseNova-U1-A3B-MoT(专家混合型30B-A3B)。从第一性原理设计,其在文本理解、视觉语言感知、知识推理、代理决策与空间智能等任务上媲美顶级理解型VLMs;同时在语义一致性和视觉保真度方面表现优异,胜任常规或知识密集型任意到图像(X2I)合成、复杂文本富集信息图生成以及带或不带思考模式的交错视觉语言生成。此外,详述模型设计、数据预处理、预训练/后训练及推理策略以支持社区研究。初步证据表明,模型在视觉语言动作(VLA)与世界模型(WM)场景中亦表现强劲,指向未来多模态智能不再依赖模态间转换,而是原生跨模态思考与行动的愿景。

原文摘要 · Abstract (English)

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.

多模态统一生成与理解视觉语言模型NEO-unify

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。