解决语音驱动虚拟人生成中动态退化问题,恢复自然表情与口型同步。
DynaForcing: Overcoming Dynamic Collapse in Self-Forcing Distillation for Streaming Avatar Generation

- 通过数据级锚定、损失级运动奖励和条件级参考扰动三策略破除动态崩溃
- 动态退化指标提升近一倍(Dyn-Deg 0.31→0.73),口型同步显著改善
- 适合追求高动态真实感的实时虚拟人应用,尤其对口型同步敏感场景
语音驱动虚拟人生成需实现逼真口型同步、丰富表情及实时流式输出。近期工作通过自强化与分布匹配蒸馏(DMD)实现实时性,但该范式存在未被系统描述的关键缺陷:动态崩溃,即学生模型收敛至近静态最优解,虽感知质量高,但时间动态严重抑制。我们溯源发现根源在于DMD中的反向KL目标偏向低运动模式,以及无锚定自条件化引发反馈环放大崩溃。尤其对虚拟人而言,微小运动损失即破坏口型同步与表情表达。为此,我们提出DynaForcing训练框架,包含三个互补策略:混合强制在数据层锚定轨迹以打破反馈环;动态感知奖励正则化通过强化学习视角引入显式运动奖励,抵消反向KL偏差;参考扰动在条件层打乱参考图像,迫使模型依赖音频驱动动作。此外引入计算图剪枝与梯度回放,使自强化推理的显存占用降低一个数量级以上。实验表明,DynaForcing将动态退化指标恢复至接近教师水平(Dyn-Deg: 0.31 → 0.73,Sync-C: 7.03 → 7.68),同时提升视觉质量,全程消除质量-动态权衡,无需早停。
原文摘要 · Abstract (English)
Audio-driven avatar generation requires realistic lip-sync, expressive motion, and real-time streaming. Recent work achieves the latter via self-forcing with Distribution Matching Distillation (DMD), but this paradigm suffers from a critical failure that has not been systematically characterized: dynamic collapse, where the student model converges to a near-static optimum with high perceptual quality but severely suppressed temporal dynamics. We trace this to two causes: the reverse KL objective in DMD, which biases toward low-motion modes, and unanchored self-conditioning, which creates a feedback loop that amplifies collapse. This is especially harmful for avatars, where even subtle motion loss breaks lip-sync and expression. To address this, we propose DynaForcing, a training framework with three complementary strategies applied at different levels. Specifically, Hybrid Forcing anchors rollouts to ground-truth dynamics at the data level to break the feedback loop. Dynamics-Aware Reward Regularization introduces explicit motion rewards via the RL interpretation of DMD to counteract the reverse KL bias at the loss level. Reference Perturbation perturbs reference images to decouple identity from static details, forcing the model to rely on audio for motion at the conditioning level. We further introduce computation graph pruning and gradient replay, reducing the GPU footprint of self-forcing by over an order of magnitude. Experiments show that DynaForcing recovers dynamics to teacher-comparable levels (Dyn-Deg: 0.31 -> 0.73, Sync-C: 7.03 -> 7.68) while improving visual quality, resolving the quality-dynamics trade-off throughout training without early stopping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。