让快速视频生成模型也能精准控制动作,靠的是轻量级推理时纠错。
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
- 用教师模型在推理时动态修正学生模型的去噪轨迹。
- 在多个蒸馏模型上显著提升动作保真度,且不损失生成速度。
- 适合需要高效又精准动作控制的视频生成应用。
无训练运动定制通过测试时计算,将参考视频的动作模式施加到视频生成器上。现有方法多针对完整的扩散模型,需大量去噪步骤,计算成本高。随着高效蒸馏模型兴起,一个自然问题浮现:能否直接将测试时运动定制应用于蒸馏生成器以利用其加速优势?然而分析表明,现有无训练技术在蒸馏模型上失效。蒸馏从根本上改变了去噪动态,且蒸馏模型的大步去噪会丢弃分数引导所需密集中间状态,导致原有运动控制策略不兼容。为此,我们提出MotionEcho,一种新型无训练测试时蒸馏框架,实现对蒸馏视频生成器的运动定制。核心思想是推理时有限使用高质量扩散教师模型,通过重去噪学生模型终点至教师的密集轨迹,形成对齐动作的干净终点,并与学生结果插值融合;同时采用自适应调度机制决定何时及如何引入教师引导。结果表明,MotionEcho通过轻量、自适应的测试时教师指导,恢复了蒸馏生成器的生成轨迹,实现准确运动控制而不牺牲生成效率。在多个蒸馏视频生成模型上的实验表明,该方法显著提升动作保真度和视觉质量,同时保持蒸馏生成的效率优势。
原文摘要 · Abstract (English)
Training-free motion customization imposes motion patterns from reference videos onto video generators through test-time computation. Most existing methods target full diffusion models, requiring many denoising steps and high computational cost. With the rise of efficient distilled models, a natural question arises: can test-time motion customization be applied directly to distilled generators with their accelerated sampling and efficiency gains? However, our analysis reveals that existing training-free techniques fail on distilled models. Distillation fundamentally alters the denoising dynamics that prior test-time guidance relies on, and the large denoising steps of distilled generators discard the dense intermediate states that score guidance requires, rendering existing motion control strategies incompatible with fast generation. To address this limitation, we propose MotionEcho, a novel training-free test-time distillation framework that enables motion customization for distilled video generators. The key idea is to correct the student model's sampling trajectory with restricted usage of a high-quality diffusion teacher at inference time. Teacher supervises the student's denoising by re-noising the student's endpoint onto its dense trajectory to form a motion-aligned clean endpoint, then interpolating it with the student's, while an adaptive scheduling mechanism determines when and how much teacher guidance is needed. As a result, MotionEcho restores generative trajectories for distilled video generators via lightweight, adaptive test-time teacher guidance, enabling accurate motion control without compromising generation efficiency. Extensive experiments on multiple distilled video generation models demonstrate that our method significantly improves motion fidelity and visual quality while retaining the efficiency advantages of distilled generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。