arXiv:2511.23199cs.CVcs.AI2025-11被引 3

大模型版布朗桥模型,实现高效图像视频生成与编辑

Vision Bridge Transformer at Scale

  • 用布朗桥直接建模输入到输出的轨迹,跳过传统去噪步骤
  • 200亿参数模型在图像视频翻译任务中表现优异
  • 适合需要精准指令控制的图像编辑与复杂视频转换场景

我们提出视觉桥接变换器(ViBT),一种大规模布朗桥模型,用于条件生成。与传统扩散模型从噪声生成数据不同,桥接模型直接建模输入与输出之间的轨迹,形成高效的数据到数据转换范式。通过将模型规模扩展至200亿和13亿参数,我们验证了其在图像和视频翻译任务中的有效性。为支持该规模,采用Transformer架构,并提出方差稳定的速度匹配目标以实现稳健训练。这些进展凸显了扩大桥接模型在基于指令的图像编辑和复杂视频翻译中的潜力。

原文摘要 · Abstract (English)

We introduce Vision Bridge Transformer (ViBT), a large-scale instantiation of Brownian Bridge Models designed for conditional generation. Unlike traditional diffusion models that transform noise into data, Bridge Models directly model the trajectory between inputs and outputs, creating an efficient data-to-data translation paradigm. By scaling these models to 20B and 1.3B parameters, we demonstrate their effectiveness for image and video translation tasks. To support this scale, we adopt a Transformer architecture and propose a variance-stabilized velocity-matching objective for robust training. Together, these advances highlight the power of scaling Bridge Models for instruction-based image editing and complex video translation.

视觉生成大模型桥接模型视频翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。