用掩码引导实现任意视频主体替换,支持个性化编辑。
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
- 通过掩码和参考图实现无特定主体限制的视频主体替换。
- 在新构建的数据集上性能超越现有方法,细节保留更佳。
- 适合需要精细视频定制的创作者或影视后期人员。
随着视频生成技术的快速发展,个性化视频编辑需求激增,其中主体替换是关键环节但研究仍不充分。现有方法多局限于特定领域(如人体动作或手物交互),或依赖间接编辑方式与模糊文本提示,影响最终质量。本文提出 DreamSwapV,一种掩码引导、主体无关、端到端的框架,可基于用户指定的掩码与参考图像,在任意视频中替换任意主体。为实现细粒度控制,引入多种条件并设计专用条件融合模块以高效整合。此外,采用自适应掩码策略,适应不同尺度与属性的主体,增强替换主体与上下文的互动。通过精心设计的两阶段数据集构建与训练方案,DreamSwapV 在 VBench 指标及首个提出的 DreamSwapV-Benchmark 上均显著优于现有方法。
原文摘要 · Abstract (English)
With the rapid progress of video generation, demand for customized video editing is surging, where subject swapping constitutes a key component yet remains under-explored. Prevailing swapping approaches either specialize in narrow domains--such as human-body animation or hand-object interaction--or rely on some indirect editing paradigm or ambiguous text prompts that compromise final fidelity. In this paper, we propose DreamSwapV, a mask-guided, subject-agnostic, end-to-end framework that swaps any subject in any video for customization with a user-specified mask and reference image. To inject fine-grained guidance, we introduce multiple conditions and a dedicated condition fusion module that integrates them efficiently. In addition, an adaptive mask strategy is designed to accommodate subjects of varying scales and attributes, further improving interactions between the swapped subject and its surrounding context. Through our elaborate two-phase dataset construction and training scheme, our DreamSwapV outperforms existing methods, as validated by comprehensive experiments on VBench indicators and our first introduced DreamSwapV-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。