arXiv:2503.12049cs.CV2025-03ICCV被引 14

用扩散模型补全视频中被遮挡的物体,让画面连贯自然。

TACO: Taming Diffusion for in-the-wild Video Amodal Completion

论文配图:TACO: Taming Diffusion for in-the-wild Video Amodal Completion
图 1 · 摘自论文原文
  • 基于预训练视频扩散模型的条件生成,利用其学到的物理一致性先验。
  • 在合成数据上渐进式微调,提升对真实复杂场景的泛化能力。
  • 适用于自动驾驶、机器人操作等场景,可支持后续重建与姿态估计任务。

人类能根据有限视觉线索推断出物体的完整形状和外观,依赖对物理世界的丰富先验知识。然而,在视频中保持多帧间的一致性完成部分可见物体仍具挑战,尤其在非结构化的野外视频中。本文针对视频非可视完成(VAC)任务,旨在给定一个指定目标物体的视觉提示后,生成整个视频中一致的完整物体。我们提出一种条件扩散模型TACO,复用预训练视频扩散模型所学的丰富且一致的特征流形。为实现对复杂野外场景的有效稳健泛化,我们系统性地在无遮挡视频上施加遮挡,构建了一个大规模合成数据集,包含多个难度级别。在此基础上,设计了渐进式微调策略:从简单恢复任务开始,逐步过渡到更复杂的场景。我们在互联网采集的多种野外视频以及自动驾驶、机器人操作和场景理解常用的不同未见数据集上验证了TACO的泛化能力。此外,还展示了TACO可用于物体重建、姿态估计等下游任务,凸显其在促进物理世界理解与推理方面的潜力。

原文摘要 · Abstract (English)

Humans can infer complete shapes and appearances of objects from limited visual cues, relying on extensive prior knowledge of the physical world. However, completing partially observable objects while ensuring consistency across video frames remains challenging for existing models, especially for unstructured, in-the-wild videos. This paper tackles the task of Video Amodal Completion (VAC), which aims to generate the complete object consistently throughout the video given a visual prompt specifying the object of interest. Leveraging the rich, consistent manifolds learned by pre-trained video diffusion models, we propose a conditional diffusion model, TACO, that repurposes these manifolds for VAC. To enable its effective and robust generalization to challenging in-the-wild scenarios, we curate a large-scale synthetic dataset with multiple difficulty levels by systematically imposing occlusions onto un-occluded videos. Building on this, we devise a progressive fine-tuning paradigm that starts with simpler recovery tasks and gradually advances to more complex ones. We demonstrate TACO's versatility on a wide range of in-the-wild videos from Internet, as well as on diverse, unseen datasets commonly used in autonomous driving, robotic manipulation, and scene understanding. Moreover, we show that TACO can be effectively applied to various downstream tasks like object reconstruction and pose estimation, highlighting its potential to facilitate physical world understanding and reasoning. Our project page is available at https://jason-aplp.github.io/TACO.

视频补全扩散模型物体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。