无需掩码即可自然插入任意参考主体,提升视频一致性与融合效果。
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
- 通过自动生成跨对数据构建新管道,解决训练数据稀缺问题。
- 在真实场景中实现主体与场景平衡,插入效果优于闭源商业方案。
- 适用于影视制作、虚拟拍摄等需要高质量主体插入的场景。
基于扩散模型的视频插入技术进展显著,但现有方法依赖复杂控制信号,且难以保持主体一致性,限制了实际应用。本文聚焦无掩码视频插入任务,针对数据稀缺、主体-场景失衡和插入不协调三大挑战提出解决方案。首先设计InsertPipe数据管道,自动构建多样化跨对数据以缓解数据稀缺;在此基础上提出OmniInsert统一框架,支持单/多主体参考的无掩码插入。为维持主体-场景平衡,引入条件特异性特征注入机制,结合渐进式训练策略,使模型有效平衡多源条件注入。同时设计主体聚焦损失函数,增强主体细节表现。为进一步提升插入和谐性,提出插入偏好优化方法,模拟人类审美偏好,并引入上下文感知重写模块,实现主体与原场景无缝融合。为填补领域评估空白,构建InsertBench基准,涵盖多样场景与精心选取的主体。在InsertBench上的评估显示,OmniInsert性能超越当前最先进的闭源商业方案。代码将开源。
原文摘要 · Abstract (English)
Recent advances in video insertion based on diffusion models are impressive. However, existing methods rely on complex control signals but struggle with subject consistency, limiting their practical applicability. In this paper, we focus on the task of Mask-free Video Insertion and aim to resolve three key challenges: data scarcity, subject-scene equilibrium, and insertion harmonization. To address the data scarcity, we propose a new data pipeline InsertPipe, constructing diverse cross-pair data automatically. Building upon our data pipeline, we develop OmniInsert, a novel unified framework for mask-free video insertion from both single and multiple subject references. Specifically, to maintain subject-scene equilibrium, we introduce a simple yet effective Condition-Specific Feature Injection mechanism to distinctly inject multi-source conditions and propose a novel Progressive Training strategy that enables the model to balance feature injection from subjects and source video. Meanwhile, we design the Subject-Focused Loss to improve the detailed appearance of the subjects. To further enhance insertion harmonization, we propose an Insertive Preference Optimization methodology to optimize the model by simulating human preferences, and incorporate a Context-Aware Rephraser module during reference to seamlessly integrate the subject into the original scenes. To address the lack of a benchmark for the field, we introduce InsertBench, a comprehensive benchmark comprising diverse scenes with meticulously selected subjects. Evaluation on InsertBench indicates OmniInsert outperforms state-of-the-art closed-source commercial solutions. The code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。