通过时空语义对齐提升动态人脸配音的稳定性与画质
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
- 设计双路径对齐机制,跨域融合空间与时间语义特征
- 在多个数据集上实现更稳定、更高质量的面部动作合成
- 适合关注语音驱动人脸生成与视频编辑的研究者
现有语音驱动人脸配音方法已取得显著进展,但空间与时间域之间的语义模糊性会严重影响动态人脸生成的稳定性。本文提出时空语义对齐(STSA)方法,引入双路径对齐机制和可微语义表示。前者通过一致信息学习(CIL)模块在多尺度下最大化互信息,减小空间与时间域流形差异;后者采用概率热图作为抗模糊引导,避免微小语义抖动导致的异常动态。大量实验表明,STSA在图像质量和合成稳定性方面均优于现有方法。预训练权重与推理代码已开源:https://github.com/SCAILab-USTC/STSA。
原文摘要 · Abstract (English)
Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We argue that aligning the semantic features from spatial and temporal domains is a promising approach to stabilizing facial motion. To achieve this, we propose a Spatial-Temporal Semantic Alignment (STSA) method, which introduces a dual-path alignment mechanism and a differentiable semantic representation. The former leverages a Consistent Information Learning (CIL) module to maximize the mutual information at multiple scales, thereby reducing the manifold differences between spatial and temporal domains. The latter utilizes probabilistic heatmap as ambiguity-tolerant guidance to avoid the abnormal dynamics of the synthesized faces caused by slight semantic jittering. Extensive experimental results demonstrate the superiority of the proposed STSA, especially in terms of image quality and synthesis stability. Pre-trained weights and inference code are available at https://github.com/SCAILab-USTC/STSA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。