让视频扩散模型可编辑时间,控制运动速度和节奏。
Making Time Editable in Video Diffusion Transformers

- 用轻量模块扩展预训练模型,实现时间可控。
- 无需重训练,即可调节视频运动快慢与结构。
- 适合需要精细控制视频时序的创作者。
当前用于视频生成的扩散变换器对时间进程和时序动态的控制能力有限。本文提出一种时间控制方法,通过在预训练的DiT基础上添加轻量级时间模块,实现对运动速度和时间结构的显式编辑,而无需重构主干网络。该方法在保持原有生成先验的基础上,扩展了可控的动态范围,使用户能够灵活调整视频的时间特性,为视频生成提供了更精细的时间控制能力。
原文摘要 · Abstract (English)
Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics. We propose a temporal-control methodology that extends a pretrained DiT with explicit time editing, allowing control over motion speed and temporal structure without redesigning the backbone. Its core implementation augments the pretrained model with a lightweight temporal module, preserving the original generative prior while expanding its controllable dynamic range.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。