arXiv:2503.01115cs.CV2025-03CVPR被引 9

WeGen让多模态生成能像聊天一样迭代优化,保持一致且创意十足。

WeGen: A Unified Model for Interactive Multimodal Generation as We Chat

  • 统一生成与理解,通过对话逐步优化内容
  • 在少指令下仍能生成多样创意结果,准确保留用户满意的细节
  • 适合需要持续修改的创意设计场景,如视频、广告制作

现有多模态生成模型难以胜任设计协作角色,常因指令不明确导致输出缺乏想象力,或无法维持与参考内容的一致性。本文提出WeGen,一个统一多模态生成与理解的模型,通过对话式交互实现迭代生成。它能在指令模糊时生成高创意内容,并根据用户反馈逐步优化已有结果或融合参考内容,同时保持用户已满意部分的一致性。为此,我们构建了一个大规模数据集,从互联网视频中提取丰富物体动态,并由先进基础模型自动标注动态描述,将两者交织成序列,使WeGen学习一致性感知生成。此外,引入提示自重写机制提升生成多样性。大量实验表明,WeGen在多个视觉生成基准上达到领先性能,展现出作为友好设计协作者的巨大潜力。代码与模型将在https://github.com/hzphzp/WeGen发布。

原文摘要 · Abstract (English)

Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack the ability to maintain consistency with the provided references. In this work, we introduce WeGen, a model that unifies multimodal generation and understanding, and promotes their interplay in iterative generation. It can generate diverse results with high creativity for less detailed instructions. And it can progressively refine prior generation results or integrating specific contents from references following the instructions in its chat with users. During this process, it is capable of preserving consistency in the parts that the user is already satisfied with. To this end, we curate a large-scale dataset, extracted from Internet videos, containing rich object dynamics and auto-labeled dynamics descriptions by advanced foundation models to date. These two information are interleaved into a single sequence to enable WeGen to learn consistency-aware generation where the specified dynamics are generated while the consistency of unspecified content is preserved aligned with instructions. Besides, we introduce a prompt self-rewriting mechanism to enhance generation diversity. Extensive experiments demonstrate the effectiveness of unifying multimodal understanding and generation in WeGen and show it achieves state-of-the-art performance across various visual generation benchmarks. These also demonstrate the potential of WeGen as a user-friendly design copilot as desired. The code and models will be available at https://github.com/hzphzp/WeGen.

多模态生成对话生成设计协作者一致性保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。