用多视角先验实现更稳定的视频物体插入
Controllable Video Object Insertion via Multi-View Priors

- 通过多视角表示提升物体外观一致性
- 显著减少身份漂移和边界伪影
- 适合需要精准控制物体插入的场景
视频物体插入将用户指定的物体放入现有动态场景中。现有方法通常依赖文本或单张参考图像进行生成,导致视角变化时物体外观约束不足,常出现身份漂移、前后景层叠错误、边界伪影和时间闪烁等问题。本文提出一种融合多视角物体先验的视频物体插入框架,将2D参考图升维为多视角表示,并采用视图一致的条件生成策略,提供稳定的身份引导和自适应外观提示。引入质量感知加权机制,降低噪声或重建不佳视图的影响。进一步设计了集成感知一致性模块,促进合理的遮挡关系、清晰边界与时间连续性。实验表明,相比基线方法,本框架在视觉质量、可控性、身份一致性及前景背景融合方面均有显著提升。
原文摘要 · Abstract (English)
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequently, object appearance is underconstrained under viewpoint changes, often leading to identity drift, incorrect foreground-background layering, boundary artifacts, and temporal flickering. In this paper, we propose a video object insertion framework that incorporates multi-view object priors to address these limitations. The framework lifts a 2D reference image into a multi-view representation and uses view-consistent conditioning to provide stable identity guidance and view-adaptive appearance cues. A quality-aware weighting mechanism reduces the influence of noisy or imperfect reconstructed views. We further introduce an Integration-Aware Consistency Module that promotes plausible occlusion, clean boundaries, and temporal continuity. Experiments demonstrate that the proposed framework improves visual quality, controllability, identity consistency, and foreground-background integration for video object insertion compared to the baseline methods. Project page: https://polarisxq.github.io/MOVI/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。