arXiv:2509.21086cs.CV2025-09被引 1

通过分步分解视频的时空结构,实现精准可控的视频概念迁移。

UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition

  • 将视频拆解为前景、背景和运动流三部分,分步控制生成
  • 引入链式提示机制,分阶段引导扩散过程,提升可控性
  • 自监督预训练+新数据集,适合视频编辑与创意生成研究者

我们提出一种新型架构UniTransfer,通过渐进式空间与扩散时间步分解,实现精确可控的视频概念迁移。具体而言,空间上将视频分解为前景主体、背景与运动流三部分;在此基础上,构建双流转单流的DiT架构,支持对各组件的细粒度控制。同时设计基于随机掩码的自监督预训练策略,增强大规模无标签视频数据下的表示学习能力。受思维链推理启发,重新审视去噪扩散过程,提出链式提示(CoP)机制,将去噪过程分为三个不同粒度阶段,并利用大语言模型(LLMs)提供阶段专属指令,逐步引导生成。我们还构建了一个以动物为中心的视频数据集OpenAnimal,用于推动和评估视频概念迁移研究。大量实验表明,该方法在多种参考图像与场景下均实现高质量且可编辑的视频迁移,视觉保真度与可控性均优于现有基线。

原文摘要 · Abstract (English)

We propose a novel architecture UniTransfer, which introduces both spatial and diffusion timestep decomposition in a progressive paradigm, achieving precise and controllable video concept transfer. Specifically, in terms of spatial decomposition, we decouple videos into three key components: the foreground subject, the background, and the motion flow. Building upon this decomposed formulation, we further introduce a dual-to-single-stream DiT-based architecture for supporting fine-grained control over different components in the videos. We also introduce a self-supervised pretraining strategy based on random masking to enhance the decomposed representation learning from large-scale unlabeled video data. Inspired by the Chain-of-Thought reasoning paradigm, we further revisit the denoising diffusion process and propose a Chain-of-Prompt (CoP) mechanism to achieve the timestep decomposition. We decompose the denoising process into three stages of different granularity and leverage large language models (LLMs) for stage-specific instructions to guide the generation progressively. We also curate an animal-centric video dataset called OpenAnimal to facilitate the advancement and benchmarking of research in video concept transfer. Extensive experiments demonstrate that our method achieves high-quality and controllable video concept transfer across diverse reference images and scenes, surpassing existing baselines in both visual fidelity and editability. Web Page: https://yu-shaonian.github.io/UniTransfer-Web/

视频生成扩散模型概念迁移可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。