arXiv:2605.23891cs.CV2026-05被引 3

无需掩码实现逼真视频物体插入,解决风格差异难题。

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework

论文配图:Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework
图 1 · 摘自论文原文
  • 双流框架同步生成视频并迁移风格,闭环反馈提升稳定性。
  • 在真实场景中实现物体自然插入,风格融合度优于现有方法。
  • 适合视频编辑、影视制作等需要高质量物体合成的场景。

无掩码视频物体插入成为一项挑战性任务,需将参考物体与源视频和谐融合。然而,当参考物体与源场景存在显著风格差异时,现有方法表现不佳。为此,我们提出Smart-Insertion-V,一个端到端的双流框架,同时进行视频插入与图像风格迁移。其中,图像流同步引导视频生成过程,并引入闭环反馈机制以确保鲁棒性。为应对多条件信号融合导致的特征纠缠与风格泄露问题,设计了双世界视图RoPE,通过时空偏移区分不同信号,训练开销低。此外,引入解耦引导模块,利用视觉语言模型进行语义推理,同时保留原始时间引导信号。为缓解和谐插入任务的数据差距,提出数据清洗流程,并将发布开源数据集。实验表明,该方法可将物体插入至合理位置,实现最优的风格融合效果。

原文摘要 · Abstract (English)

Mask-free video object insertion has emerged as a challenging task, requiring harmonious integration of reference objects into source videos. However, existing methods struggle when references exhibit severe stylistic domain gaps with the source scene. To overcome this, we propose \textit{\textbf{Smart-Insertion-V}}, an end-to-end \textbf{Dual-Stream} framework that concurrently conducts video insertion and image style transfer. Within this framework, the image stream synchronously guides the video generation process, while a \textbf{Closed-loop Feedback} mechanism is further incorporated to ensure robust insertion. Inevitably, integrating these diverse conditioning signals results in feature entanglement and style leakage. To tackle this issue, we design \textbf{Dual-World-View RoPE} to distinguish different signals via spatial-temporal offsets without incurring heavy training overhead. Furthermore, to facilitate spatial grounding and stylistic adaptation, we introduce a \textbf{Decoupled Guidance Module} that leverages a Vision-Language Model for semantic reasoning while preserving original temporal guidance with native text encoder. To bridge data gap for harmonious reference insertion task, we propose a data curation pipeline and will release an \textbf{open-source dataset}. Experiments demonstrate that our method can insert objects into plausible positions while achieving the most harmonious results.

视频插入风格迁移双流框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。