通过对比掩码预训练,让音视频人脸同步更精准。
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
- 用三个提示令牌分离身份、语音驱动动作和环境动作。
- 在四个下游任务中均达顶尖性能,尤其在语音对齐上显著提升。
- 适合做音视频同步、面部表情识别、口型识别与视觉配音的开发者。
我们提出SyncLipMAE,一种自监督预训练框架,从无标签音视频流中学习同步感知且可迁移的人脸动态。该方法将掩码视觉建模与跨模态对比对齐结合,使用三个每帧提示令牌显式编码说话人脸帧的关键因素:身份、语音运动(与语音同步的面部动作)以及环境运动(如眨眼、头部姿态等非音频相关动作)。对比目标以时间对齐的语音运动与音频令牌为正样本,错位对为负样本,推动双模态进入共享嵌入空间,实现细粒度音视频流同步。预训练后,对齐的音频令牌与视觉提示令牌(身份、语音运动、环境运动)构成统一接口,支持四大不同下游任务:(i) 音视频流同步;(ii) 面部情绪与头部/面部动作识别;(iii) 视觉语音识别;(iv) 视觉配音,实现单模型内不可区分的音频或视频驱动控制。在涵盖四种任务族的多项任务中,SyncLipMAE均取得最先进性能,验证了同步感知与因子化自监督预训练的有效性。
原文摘要 · Abstract (English)
We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with cross-modal contrastive alignment and employs three per-frame prompt tokens that explicitly encode the essential factors of a talking-face frame - identity, vocal motion (speech-synchronized facial dynamics), and ambient motion (audio-agnostic movements such as blinks and head pose). The contrastive objective uses time-aligned vocal-motion and audio tokens as positives and misaligned pairs as negatives, driving both modalities into a shared embedding space and yielding token-level audio-visual stream synchronization. After pretraining, the aligned audio tokens together with the visual prompt tokens (identity, vocal motion, ambient motion) form a unified interface for four disparate downstream settings: (i) audio-visual stream synchronization; (ii) facial emotion and head/face action recognition; (iii) visual speech recognition; and (iv) visual dubbing, for which we enable indistinguishable audio- or video-driven control within a single model. Across four task families that require distinct capabilities, SyncLipMAE achieves state-of-the-art results, underscoring the effectiveness of synchronization-aware, factorized self-supervised pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。