一个模型搞定图文视频生成与编辑,靠的是混合条件输入。
VINO: A Unified Visual Generator with Interleaved OmniModal Context
- 用交错编码的多模态条件引导扩散过程,统一处理文本图像视频。
- 支持多参考物定位、长指令遵循和动态静态内容身份一致性保留。
- 适合需要跨模态生成编辑的开发者与创意工作者使用。
我们提出 VINO,一种统一的视觉生成框架,可在单一模型中完成图像与视频的生成和编辑。不同于依赖任务专用模型或各模态独立模块的做法,VINO 采用共享的扩散主干网络,以文本、图像和视频作为条件输入,实现广泛的视觉创作与编辑任务。具体而言,VINO 将视觉语言模型(VLM)与多模态扩散变换器(MMDiT)结合,将多模态输入编码为交错的条件标记,并用于指导扩散过程。该设计支持多参考物定位、长形式指令遵循以及静态与动态内容间的一致性身份保持,同时避免了模态特定的架构组件。为训练这一统一系统,我们引入多阶段训练流程,逐步将视频生成基础模型扩展为可同时处理图像与视频输入输出的多任务生成器。在多种生成与编辑基准测试中,VINO 展现出优异的视觉质量、忠实的指令遵循能力、改进的参考物与属性保持,以及更可控的多身份编辑性能。结果表明,这为可扩展的统一视觉生成提供了可行路径,也凸显了交错式上下文计算在通用视觉创作中的潜力。
原文摘要 · Abstract (English)
We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion backbone that conditions on text, images and videos, enabling a broad range of visual creation and editing tasks under one model. Specifically, VINO couples a vision-language model (VLM) with a Multimodal Diffusion Transformer (MMDiT), where multimodal inputs are encoded as interleaved conditioning tokens, and then used to guide the diffusion process. This design supports multi-reference grounding, long-form instruction following, and coherent identity preservation across static and dynamic content, while avoiding modality-specific architectural components. To train such a unified system, we introduce a multi-stage training pipeline that progressively expands a video generation base model into a unified, multi-task generator capable of both image and video input and output. Across diverse generation and editing benchmarks, VINO demonstrates strong visual quality, faithful instruction following, improved reference and attribute preservation, and more controllable multi-identity edits. Our results highlight a practical path toward scalable unified visual generation, and the promise of interleaved, in-context computation as a foundation for general-purpose visual creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。