统一视频时空迁移框架,实现更灵活高保真生成
OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
- 用多帧视角信息提升外观一致性,利用时序线索实现精细控制
- 在外观和时序迁移上超越现有方法,运动迁移媲美姿态引导模型
- 适合需要灵活视频编辑的创作者,尤其擅长风格与动态同步迁移
视频比图像或文本包含更丰富的信息,能捕捉空间与时间动态。然而,现有视频定制方法多依赖参考图像或任务特定的时间先验,未能充分利用视频固有的丰富时空信息,限制了生成的灵活性与泛化能力。为此,我们提出OmniTransfer,一个统一的时空视频迁移框架。它通过跨帧多视角信息增强外观一致性,利用时序线索实现细粒度时间控制。为统一各类视频迁移任务,OmniTransfer引入三项关键设计:任务感知位置偏置,自适应利用参考视频信息以改善时间对齐或外观一致性;参考解耦因果学习,分离参考与目标分支,实现精确参考迁移并提升效率;任务自适应多模态对齐,利用多模态语义引导动态区分并应对不同任务。大量实验表明,OmniTransfer在外观(身份与风格)和时序迁移(摄像机运动与视频效果)方面优于现有方法,且在无需姿态信息的情况下达到姿态引导方法的运动迁移水平,建立了一种灵活、高保真的视频生成新范式。
原文摘要 · Abstract (English)
Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the rich spatio-temporal information inherent in videos, thereby limiting flexibility and generalization in video generation. To address these limitations, we propose OmniTransfer, a unified framework for spatio-temporal video transfer. It leverages multi-view information across frames to enhance appearance consistency and exploits temporal cues to enable fine-grained temporal control. To unify various video transfer tasks, OmniTransfer incorporates three key designs: Task-aware Positional Bias that adaptively leverages reference video information to improve temporal alignment or appearance consistency; Reference-decoupled Causal Learning separating reference and target branches to enable precise reference transfer while improving efficiency; and Task-adaptive Multimodal Alignment using multimodal semantic guidance to dynamically distinguish and tackle different tasks. Extensive experiments show that OmniTransfer outperforms existing methods in appearance (ID and style) and temporal transfer (camera movement and video effects), while matching pose-guided methods in motion transfer without using pose, establishing a new paradigm for flexible, high-fidelity video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。