让物体插入视频时有真实光影和空间位置,效果更自然。
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
- 通过单帧设定3D姿态,自动追踪视频中物体位置与遮挡。
- 引入光学对齐机制,实现阴影反射等物理光照效果延伸。
- 自建数据集ROSE++,支持光影效果的监督学习,适合影视制作。
扩散模型虽已推动视频编辑发展,但高质量视频物体插入(VOI)仍受限于缺乏4D场景理解及正确的光学交互(如阴影、反射)。为此,我们提出InsertAnywhere框架,实现几何精准定位与光学感知的视频合成。首先,利用4D感知掩码生成模块,在单帧中锚定物体3D姿态,并自动传播至整个视频,准确处理局部动态与遮挡。为实现真实的光照交互,引入光学感知表示对齐策略,通过扩展掩码引导特征提取,使阴影、反射等效果自然延伸至物体边界之外。针对训练数据缺失问题,构建并开源了专用于光学效应监督学习的四元组数据集ROSE++。大量实验表明,InsertAnywhere在复杂真实场景中生成的插入结果兼具几何合理性与光度真实性,显著优于现有研究及商业生成工具。
原文摘要 · Abstract (English)
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D scene understanding and a lack of proper optical interactions, such as shadows and reflections. To address these limitations, we present InsertAnywhere, a comprehensive VOI framework that achieves geometrically grounded object placement and optics-aware video synthesis. Our approach first leverages a 4D-aware mask generation module that allows users to anchor an object's 3D pose in a single frame. The framework automatically propagates this placement across the video, accurately handling local scene dynamics and occlusions. To synthesize realistic physical lighting interactions, we introduce Optics-Aware Representation Alignment, a novel strategy that utilizes an extended mask to guide feature extraction, enabling optical effects to seamlessly extend beyond the inserted object's boundary. Finally, to overcome the lack of training data for such phenomena, we construct and open-source ROSE++, a specialized quadruplet dataset tailored for the supervised learning of optical effects. Extensive experiments demonstrate that InsertAnywhere produces geometrically plausible and photometrically realistic insertions in complex real-world scenarios, significantly outperforming existing research and commercial generative tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。