让AI同时生成文字和图像并保持一致性。
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
- 分两步训练:先规划文本,再根据文本生成图像。
- 仅用文本模拟的图像数据,就能生成长文本连贯且图像一致的内容。
- 适合需要图文交替生成的创作任务,如故事插画、多模态对话。
近年来统一模型在理解和生成方面取得显著进展,但多数模型虽能接受多模态输入,却只能输出单模态内容。这一问题主要源于训练数据稀缺及长距离跨模态上下文建模困难。为此,我们提出将交织生成分解为文本规划与视觉一致性建模,并构建包含规划器(planner)和可视化器(visualizer)的框架。规划器生成视觉内容的密集文本描述,可视化器据此合成图像。在此引导下,我们构建大规模文本代理交织数据(以文本表示视觉内容)用于训练规划器,并收集参考引导的图像数据训练可视化器。该设计使Wan-Weaver展现出涌现的交织生成能力,具备长程文本连贯性与视觉一致性。同时,融合多样化理解与生成数据进行规划器训练,使模型具备强任务推理与生成能力。为评估模型在交织生成中的表现,我们进一步构建涵盖多维度应用场景的基准测试。大量实验表明,即使未接触任何真实交织数据,Wan-Weaver仍优于现有方法。
原文摘要 · Abstract (English)
Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of producing interleaved content is mainly due to training data scarcity and the difficulty of modeling long-range cross-modal context. To address this issue, we decompose interleaved generation into textual planning and visual consistency modeling, and introduce a framework consisting of a planner and a visualizer. The planner produces dense textual descriptions for visual content, while the visualizer synthesizes images accordingly. Under this guidance, we construct large-scale textual-proxy interleaved data (where visual content is represented in text) to train the planner, and curate reference-guided image data to train the visualizer. These designs give rise to Wan-Weaver, which exhibits emergent interleaved generation ability with long-range textual coherence and visual consistency. Meanwhile, the integration of diverse understanding and generation data into planner training enables Wan-Weaver to achieve robust task reasoning and generation proficiency. To assess the model's capability in interleaved generation, we further construct a benchmark that spans a wide range of use cases across multiple dimensions. Extensive experiments demonstrate that, even without access to any real interleaved data, Wan-Weaver achieves superior performance over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。