揭秘文本生成视频中运动与身份的纠缠,实现高效零样本动作迁移。
Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation
- 通过查询特征控制运动与身份的耦合关系
- 零样本动作迁移效率提升10倍,且保持身份一致
- 无需训练即可实现多镜头角色身份统一
文本到视频的扩散模型在从文本描述生成连贯视频方面取得了显著进展。然而,这些模型中运动、结构与身份表征之间的相互作用仍缺乏深入研究。本文分析发现,自注意力查询(Q)特征同时影响布局、运动和主体身份,在去噪过程中对身份有显著影响,导致动作迁移时难以避免身份转移。基于此,我们提出了查询特征注入(Q注入)的控制方法,并实现了两个应用:(1) 基于VideoCrafter2和WAN 2.1的零样本动作迁移,效率比现有方法高10倍;(2) 一种无需训练的一致性多镜头视频生成技术,角色在多个镜头间保持身份一致,同时通过Q注入提升动作保真度。
原文摘要 · Abstract (English)
Text-to-video diffusion models have shown remarkable progress in generating coherent video clips from textual descriptions. However, the interplay between motion, structure, and identity representations in these models remains under-explored. Here, we investigate how self-attention query (Q) features simultaneously govern motion, structure, and identity and examine the challenges arising when these representations interact. Our analysis reveals that Q affects not only layout, but that during denoising Q also has a strong effect on subject identity, making it hard to transfer motion without the side-effect of transferring identity. Understanding this dual role enabled us to control query feature injection (Q injection) and demonstrate two applications: (1) a zero-shot motion transfer method - implemented with VideoCrafter2 and WAN 2.1 - that is 10 times more efficient than existing approaches, and (2) a training-free technique for consistent multi-shot video generation, where characters maintain identity across multiple video shots while Q injection enhances motion fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。