让说话头像生成在推理时自动调整身份特征,解决动态表情与固定图像不匹配问题。
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

- 推理时通过自反馈循环重构参考图像,动态适配面部运动变化。
- 在多个数据集上提升唇音同步率、时间连贯性及身份一致性,显著改善画质。
- 无需训练或微调,适用于任意预训练生成模型,适合视频动画开发者使用。
音频驱动的说话头像生成虽已取得显著进展,但多数方法依赖单一静态参考图像作为整个视频生成过程的条件,导致固定身份特征与动态面部运动之间出现不匹配,引发身份漂移、时间不一致和感知质量下降。本文提出测试时自适应条件化(TT-SAC),一种无参数的推理框架,使预训练说话头像生成器能在不重新训练、不梯度更新、无额外监督的情况下,在推理阶段自适应调整其条件表示。该方法将生成器输出重新编码,形成更契合合成序列时序动态的优化条件表示,单次适应即逼近生成过程的自洽平衡,稳定身份与动作。理论分析表明,在弱Lipschitz假设下,该方法降低特征方差,提升生成稳定性,并呈现可解释的偏差-方差权衡机制。在先进生成器与基准数据集上的大量实验验证了其在唇音同步精度、时间连贯性、身份保持与感知保真度上的持续提升。TT-SAC为生成视频模型提供了一种模型无关、训练零成本的增强策略,确立了测试时条件适应作为稳定音频驱动肖像动画的有效机制。
原文摘要 · Abstract (English)
Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, most existing approaches rely on a single static reference image to condition the entire video generation process at inference stage. This static conditioning paradigm often creates a mismatch between fixed identity features and dynamically evolving facial motion, leading to identity drift, temporal inconsistency, and degraded perceptual quality. We introduce Test-Time Self-Adaptive Conditioning (TT-SAC), a parameter-free inference framework that enables pretrained talking-head generators to adapt their conditioning representations during inference without retraining, gradient updates, or additional supervision. Instead of treating the reference portrait as immutable, TT-SAC composes the generator with its encoder in a feedback loop: the generator's own outputs are re-encoded to construct a refined conditioning representation that better aligns with the temporal dynamics of the synthesized sequence. A single adaptation step approximates a self-consistent equilibrium of the generative process, stabilizing identity and motion across time. We further provide theoretical analysis showing that test-time conditioning adaptation reduces feature variance and improves generative stability under mild Lipschitz assumptions, while exhibiting a principled bias-variance tradeoff that governs the optimal strength of adaptation. Extensive experiments on state-of-the-art talking-head generators and benchmark datasets demonstrate consistent improvements in lip-sync accuracy, temporal coherence, identity preservation, and perceptual fidelity. TT-SAC offers a model-agnostic and training-free strategy for enhancing generative video models, establishing test-time conditioning adaptation as an effective mechanism for stabilizing audio-driven portrait animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。