用扩散模型实现任意主体视频虚化,支持可控焦点与平滑过渡。
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
- 基于多平面图像的聚焦面条件控制,利用预训练模型3D先验。
- 在合成与真实数据上实现更连贯的时序效果和精确的空间模糊。
- 适合需要高质量可控景深特效的视频编辑与虚拟拍摄场景。
扩散模型已成为模拟相机效果的强大工具,可实现几何变换与逼真光学效果。其中,基于图像的虚化渲染已取得良好进展,但视频虚化仍缺乏探索。现有图像方法存在时间闪烁与模糊过渡不一致问题,而当前视频编辑方法对焦平面与虚化强度缺乏显式控制,限制了可控视频虚化的应用。本文提出一种一步式扩散框架,实现时序一致、深度感知的视频虚化渲染。该框架采用适配焦平面的多平面图像(MPI)表示作为视频扩散模型的条件,从而利用预训练主干网络中的强大3D先验。为进一步提升时序稳定性、深度鲁棒性与细节保留,引入渐进式训练策略。在合成与真实世界基准上的实验表明,本方法在时序连贯性、空间精度与可控性方面均优于先前基线。这是首个专门用于视频虚化生成的扩散框架,为时序一致且可控的景深效果建立了新基准。
原文摘要 · Abstract (English)
Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based bokeh rendering has shown promising results, but diffusion for video bokeh remains unexplored. Existing image-based methods are plagued by temporal flickering and inconsistent blur transitions, while current video editing methods lack explicit control over the focus plane and bokeh intensity. These issues limit their applicability for controllable video bokeh. In this work, we propose a one-step diffusion framework for generating temporally coherent, depth-aware video bokeh rendering. The framework employs a multi-plane image (MPI) representation adapted to the focal plane to condition the video diffusion model, thereby enabling it to exploit strong 3D priors from pretrained backbones. To further enhance temporal stability, depth robustness, and detail preservation, we introduce a progressive training strategy. Experiments on synthetic and real-world benchmarks demonstrate superior temporal coherence, spatial accuracy, and controllability, outperforming prior baselines. This work represents the first dedicated diffusion framework for video bokeh generation, establishing a new baseline for temporally coherent and controllable depth-of-field effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。