无需训练即可实现物体在视频中自然插入,保持背景一致且交互真实。
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

- 分步编辑+语义运动描述,解耦视频物体插入任务。
- 相比顶尖方法,PSNR提升18.8%,SSIM提升20.1%,LPIPS降低44.1%。
- 适合需要快速高保真视频编辑的创作者与影视制作团队。
视频物体插入需保证时空一致性与交互真实性,远超简单内容放置。现有方法常依赖显式运动设计或资源密集型重训练,限制灵活性与泛化能力。为此,我们提出无需训练的SimInsert,将任务解耦为直观的单帧编辑与语义运动描述。利用图像到视频扩散模型的强生成先验,SimInsert实现编辑的时序传播,在严格保持背景不变的同时,支持文本驱动的插入物体与动态环境间合理交互。方法基于非侵入式引导机制,确保结构一致性,促进无缝边界融合,并抑制去噪轨迹中的保真度漂移。大量定量实验验证其有效性:相比现有最优方法,PSNR提升18.8%,SSIM提升20.1%,LPIPS下降44.1%,提供高效高保真视频编辑方案。
原文摘要 · Abstract (English)
Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or resource-intensive retraining, restricting their flexibility and generalization. To bridge this gap, we present \textit{SimInsert}, a training-free paradigm that efficiently decouples the task into intuitive single-frame editing and semantic motion description. By harnessing the robust generative priors of image-to-video diffusion models, SimInsert propagates edits temporally, strictly preserving background invariance while enabling plausible, text-driven interactions between the inserted object and the dynamic environment. Our approach hinges on non-invasive guidance mechanisms that enforce structural consistency, facilitate seamless boundary fusion, and counteract the fidelity drift that typically accumulates during the denoising trajectory. Extensive quantitative experiments validate our efficacy: SimInsert surpasses state-of-the-art methods with an 18.8\% gain in PSNR, 20.1\% in SSIM, and a 44.1\% decrease in LPIPS, offering a streamlined solution for high-fidelity video editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。