揭示视频扩散模型中运动与外观在时间步的解耦规律,可提升动作迁移效果。
Characterizing Motion Encoding in Video Diffusion Timesteps
- 通过条件注入实验量化运动与外观在不同时间步的竞争关系。
- 发现早期时间步主导运动,后期时间步主导外观,存在明确解耦边界。
- 仅在早期阶段训练和推理,即可实现强动作迁移,无需额外模块。
文本到视频的扩散模型通过迭代去噪生成时空运动与视觉外观,但运动在时间步中的编码机制尚不明确。实践中常依赖经验法则:早期时间步主要决定运动与布局,后期则细化外观,但该现象缺乏系统分析。本文通过在特定时间步范围注入新条件,以外观编辑与运动保持之间的权衡作为运动编码的代理指标,开展大规模定量研究。该方法能定量刻画运动与外观在去噪轨迹上的竞争关系,揭示在多种架构下均存在早期运动主导、后期外观主导的稳定模式,从而确定时间步空间中的运动-外观解耦边界。基于此,我们简化了单次动作定制范式,仅将训练与推理限制在运动主导区间,无需辅助去偏模块或特殊目标函数,即可实现优异的动作迁移效果。本分析将常用经验法则转化为时空解耦原则,所提时间步约束方案可直接集成至现有动作迁移与编辑方法中。
原文摘要 · Abstract (English)
Text-to-video diffusion models synthesize temporal motion and spatial appearance through iterative denoising, yet how motion is encoded across timesteps remains poorly understood. Practitioners often exploit the empirical heuristic that early timesteps mainly shape motion and layout while later ones refine appearance, but this behavior has not been systematically characterized. In this work, we proxy motion encoding in video diffusion timesteps by the trade-off between appearance editing and motion preservation induced when injecting new conditions over specified timestep ranges, and characterize this proxy through a large-scale quantitative study. This protocol allows us to factor motion from appearance by quantitatively mapping how they compete along the denoising trajectory. Across diverse architectures, we consistently identify an early, motion-dominant regime and a later, appearance-dominant regime, yielding an operational motion-appearance boundary in timestep space. Building on this characterization, we simplify current one-shot motion customization paradigm by restricting training and inference to the motion-dominant regime, achieving strong motion transfer without auxiliary debiasing modules or specialized objectives. Our analysis turns a widely used heuristic into a spatiotemporal disentanglement principle, and our timestep-constrained recipe can serve as ready integration into existing motion transfer and editing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。