arXiv:2411.16683cs.CV2024-11CVPR被引 39

无需固定背景,可生成动态遮挡区域的视频分层分解方法

Generative Omnimatte: Learning to Decompose Video into Layers

  • 用扩散模型学习移除特定物体带来的场景效应
  • 在无姿态/深度信息下仍能完整生成遮挡区域
  • 适合处理含软阴影、反光、溅水等复杂视频

给定一段视频和一组目标物体掩码,全景蒙版方法旨在将视频分解为语义明确的图层,包含各个物体及其伴随效果(如阴影、反射)。现有方法依赖静态背景或精确的相机位姿与深度估计,在这些假设不成立时表现不佳。此外,由于缺乏自然视频的生成先验,现有方法无法补全动态遮挡区域。本文提出一种新型生成式分层视频分解框架,无需假设场景静止,也无需相机位姿或深度信息,可生成干净完整的图层,包括对动态遮挡区域的合理补全。核心思想是训练一个视频扩散模型,识别并移除特定物体引起的场景效应。我们证明该模型可通过少量精心构建的数据集,从现有视频修复模型微调获得,并在大量随意拍摄的视频上实现高质量分解与编辑结果,涵盖软阴影、镜面反射、飞溅水流等多种复杂场景。

原文摘要 · Abstract (English)

Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections. Existing omnimatte methods assume a static background or accurate pose and depth estimation and produce poor decompositions when these assumptions are violated. Furthermore, due to the lack of generative prior on natural videos, existing methods cannot complete dynamic occluded regions. We present a novel generative layered video decomposition framework to address the omnimatte problem. Our method does not assume a stationary scene or require camera pose or depth information and produces clean, complete layers, including convincing completions of occluded dynamic regions. Our core idea is to train a video diffusion model to identify and remove scene effects caused by a specific object. We show that this model can be finetuned from an existing video inpainting model with a small, carefully curated dataset, and demonstrate high-quality decompositions and editing results for a wide range of casually captured videos containing soft shadows, glossy reflections, splashing water, and more.

视频分解扩散模型图像修复动态遮挡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。