用统一视觉流实现多模态生成,输入变图像,输出更精准。
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- 将文本、布局、编辑指令全转为视觉提示,实现图像输入-输出的统一流程
- 在500万视觉提示数据上训练,跨任务性能领先开源模型
- 适合追求端到端视觉生成、减少模态对齐复杂性的研究者和开发者
多模态生成长期依赖以文本驱动的流水线,语言虽能指导视觉却无法在其中推理或创造。我们挑战这一范式,提出所有模态(包括文本描述、空间布局和编辑指令)能否统一为单一视觉表征。为此,我们提出FlowInOne框架,将多模态生成重构为纯视觉流,将所有输入转化为视觉提示,实现由单一流匹配模型驱动的图像输入-图像输出管道。该视觉中心化设计天然消除了跨模态对齐瓶颈、噪声调度及任务专用结构分支,统一了文本到图像生成、布局引导编辑和视觉指令遵循。为支持该框架,我们构建了包含500万视觉提示对的VisPrompt-5M数据集,涵盖物理感知力动态与轨迹预测等多样任务,并推出VP-Bench基准,严格评估指令忠实性、空间精度、视觉真实性和内容一致性。大量实验表明,FlowInOne在所有统一生成任务中均达到开源模型最优性能,且与顶尖商业系统相当,确立了完全视觉中心化生成建模的新范式,使感知与创造共存于统一连续视觉空间中。代码与模型已公开。
原文摘要 · Abstract (English)
Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions, spatial layouts, and editing instructions, can be unified into a single visual representation. We present FlowInOne, a framework that reformulates multimodal generation as a purely visual flow, converting all inputs into visual prompts and enabling a clean image-in, image-out pipeline governed by a single flow matching model. This vision-centric formulation naturally eliminates cross-modal alignment bottlenecks, noise scheduling, and task-specific architectural branches, unifying text-to-image generation, layout-guided editing, and visual instruction following under one coherent paradigm. To support this, we introduce VisPrompt-5M, a large-scale dataset of 5 million visual prompt pairs spanning diverse tasks including physics-aware force dynamics and trajectory prediction, alongside VP-Bench, a rigorously curated benchmark assessing instruction faithfulness, spatial precision, visual realism, and content consistency. Extensive experiments demonstrate that FlowInOne achieves state-of-the-art performance among open-source models across all unified generation tasks while remaining competitive with leading commercial systems, thereby establishing a new foundation for fully vision-centric generative modeling, in which perception and creation coexist within a unified continuous visual space. Our code and models are released on https://csu-jpg.github.io/FlowInOne.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。