提出时序对齐评估方法,让语音驱动人脸生成更真实地反映自然说话节奏。
Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

- 用软动态时间规整实现序列级对齐,避免帧间严格匹配的误差
- 在7个数据集上测试20种方法,验证对齐评估更稳定、更一致
- 适合关注生成质量与自然度平衡的研究者和开发者
语音驱动人脸生成技术发展迅速,但现有评估多依赖帧级指标,假设生成视频与参考视频严格时序对齐。然而,语音驱动的面部运动天然存在轻微时间偏移、语速差异和风格变化,传统指标会将这些合理差异误判为质量问题,导致方法比较不公平。本文主张将动态生成模型的评估视为序列对齐问题而非独立帧对比。我们引入统一的序列级重构框架,将软动态时间规整(Soft DTW)集成到现有评估流程中,通过保持时间顺序对齐特征轨迹,在不改变原有感知、身份或同步编码器的前提下,提升对有限时序错位的鲁棒性。帧级评估可看作刚性对齐的特例;而序列级对齐展现出更强稳定性、更低对时间偏差的敏感性,并更清晰区分不同建模范式。基于此,我们在七大数据集(涵盖标准、野外、风格多样场景)上对20种方法进行大规模基准测试,结果表明:时序对齐指标对时间差异更鲁棒,跨数据集结果更一致,能更好揭示建模范式间的系统性权衡,如同步性与真实性、表现力与稳定性之间的取舍。
原文摘要 · Abstract (English)
Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos. This assumption does not match speech-driven facial motion, which naturally includes slight timing shifts, different speaking speeds, and stylistic variations. As a result, conventional metrics may treat harmless timing differences as quality errors, making it harder to fairly compare methods and understand their trade-offs. In this work, we argue that evaluation of dynamic generative models should be formulated as a sequence-alignment problem rather than independent frame comparison. We introduce a unified sequence-level reformulation that integrates Soft Dynamic Time Warping into established evaluation pipelines. By aligning feature trajectories while preserving temporal order, the proposed framework provides robustness to bounded temporal misalignments without altering the underlying perceptual, identity, or synchronization encoders. We show that frame-wise evaluation can be viewed as a special case under rigid alignment, while sequence-level alignment provides improved stability, lower sensitivity to timing differences, and clearer separation between modeling paradigms. Building on this principled formulation, we conduct a large-scale benchmark of 20 methods across seven datasets spanning canonical, in-the-wild, and style-diverse scenarios under standardized protocols. Extensive experiments show that temporally aligned metrics are more robust to timing differences, provide more consistent results across datasets, and better reveal systematic trade-offs between modeling paradigms, such as synchronization versus realism and expressiveness versus stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。