arXiv:2605.04128cs.GRcs.AI2026-05被引 4

统一视觉理解与生成,让模型具备空间智能。

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

论文配图:JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
图 1 · 摘自论文原文
  • 用增强空间感知的多模态大模型+扩散生成结构,实现感知与生成协同。
  • 在理解、生成、长文本渲染和编辑任务上均达到顶尖表现。
  • 适合构建具空间推理能力的视觉-语言-行动系统。

我们提出JoyAI-Image,一个统一的多模态基础模型,支持视觉理解、文本到图像生成及指令引导的图像编辑。该模型结合空间增强的多模态大语言模型(MLLM)与多模态扩散变换器(MMDiT),通过共享的多模态接口实现感知与生成的双向交互。在此架构基础上,我们设计了一套可扩展的训练方案,融合统一指令微调、长文本渲染监督、空间对齐数据以及通用与空间编辑信号。这一设计使模型兼具广泛多模态能力,并强化几何感知推理与可控视觉合成。在理解、生成、长文本渲染和编辑等基准测试中,JoyAI-Image表现出当前最优或极具竞争力的性能。更重要的是,增强理解、可控空间编辑与新视角辅助推理之间的双向循环,推动模型从一般视觉能力迈向更强的空间智能。这些结果为视觉-语言-行动系统和世界模型等下游应用提供了可行路径。

原文摘要 · Abstract (English)

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language Model (MLLM) with a Multimodal Diffusion Transformer (MMDiT), allowing perception and generation to interact through a shared multimodal interface. Around this architecture, we build a scalable training recipe that combines unified instruction tuning, long-text rendering supervision, spatially grounded data, and both general and spatial editing signals. This design gives the model broad multimodal capability while strengthening geometry-aware reasoning and controllable visual synthesis. Experiments across understanding, generation, long-text rendering, and editing benchmarks show that JoyAI-Image achieves state-of-the-art or highly competitive performance. More importantly, the bidirectional loop between enhanced understanding, controllable spatial editing, and novel-view-assisted reasoning enables the model to move beyond general visual competence toward stronger spatial intelligence. These results suggest a promising path for unified visual models in downstream applications such as vision-language-action systems and world models.

多模态空间智能图像生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。