arXiv:2512.13690cs.CVcs.AI2025-12

让视频生成过程可交互预览,实时查看中间结果并调控生成。

DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders

  • 通过多分支解码器在任意步骤生成实时预览
  • 4秒视频预览耗时小于1秒,速度超4倍实时
  • 支持噪声步调控与模态引导,适合交互式创作

视频扩散模型虽推动了生成视频的发展,但其生成过程精度低、速度慢且不透明,用户需长时间等待。本文提出 DiffusionBrowser,一种模型无关、轻量级的解码框架,可在去噪过程中的任意时间步或变换器块处实现交互式预览。该模型能以超过4倍实时速度(4秒视频少于1秒)生成包含RGB和场景内在信息的多模态预览,保持与最终视频一致的外观和运动。借助训练好的解码器,我们实现了通过随机性重注入和模态引导在中间噪声步骤交互控制生成,解锁新的可控能力。此外,我们系统性地利用学习到的解码器探测模型,揭示了场景、物体等细节在原本黑箱的去噪过程中如何被组合与构建。

原文摘要 · Abstract (English)

Video diffusion models have revolutionized generative video synthesis, but they are imprecise, slow, and can be opaque during generation -- keeping users in the dark for a prolonged period. In this work, we propose DiffusionBrowser, a model-agnostic, lightweight decoder framework that allows users to interactively generate previews at any point (timestep or transformer block) during the denoising process. Our model can generate multi-modal preview representations that include RGB and scene intrinsics at more than 4$\times$ real-time speed (less than 1 second for a 4-second video) that convey consistent appearance and motion to the final video. With the trained decoder, we show that it is possible to interactively guide the generation at intermediate noise steps via stochasticity reinjection and modal steering, unlocking a new control capability. Moreover, we systematically probe the model using the learned decoders, revealing how scene, object, and other details are composed and assembled during the otherwise black-box denoising process.

视频生成扩散模型交互式生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。