arXiv:2601.12066cs.CV2026-01中稿 · ICML被引 1

用视频到视频的随机桥模型,更精准地移除视频物体。

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

  • 构建从含物视频到去物视频的随机路径,利用原视频结构引导生成。
  • 在多个数据集上实现更高画质和更强时序一致性,尤其适合大物体移除。
  • 自适应掩码调制机制,平衡背景保真与生成灵活性,适合实际应用。

现有视频物体移除方法多基于扩散模型的噪声到数据范式,从无信息的高斯噪声开始生成,忽略了原始视频中的丰富结构与上下文先验,导致缺乏足够指导,常出现物体清除不全或生成不符合物理逻辑的内容。本文将视频物体移除重新定义为通过随机桥模型实现的视频到视频翻译任务。与噪声初始化方法不同,本框架建立从源视频(含物体)到目标视频(去物)的直接随机路径,有效利用输入视频作为强结构先验,指导模型精确移除物体并确保填充区域与环境逻辑一致。针对强桥先验阻碍大物体移除的问题,提出新颖的自适应掩码调制策略,动态根据掩码特性调节输入嵌入,在背景保真与生成灵活性间取得平衡。大量实验表明,该方法在视觉质量与时序一致性上显著优于现有方法。

原文摘要 · Abstract (English)

Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This approach discards the rich structural and contextual priors present in the original input video. Consequently, such methods often lack sufficient guidance, leading to incomplete object erasure or the synthesis of implausible content that conflicts with the scene's physical logic. In this paper, we reformulate video object removal as a video-to-video translation task via a stochastic bridge model. Unlike noise-initialized methods, our framework establishes a direct stochastic path from the source video (with objects) to the target video (objects removed). This bridge formulation effectively leverages the input video as a strong structural prior, guiding the model to perform precise removal while ensuring that the filled regions are logically consistent with the surrounding environment. To address the trade-off where strong bridge priors hinder the removal of large objects, we propose a novel adaptive mask modulation strategy. This mechanism dynamically modulates input embeddings based on mask characteristics, balancing background fidelity with generative flexibility. Extensive experiments demonstrate that our approach significantly outperforms existing methods in both visual quality and temporal consistency. The project page is https://bridgeremoval.github.io/.

视频修复生成模型随机桥

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。