arXiv:2601.02753eess.AS2026-01

用人脸匹配声音,让合成语音更符合参考人脸特征。

Vclip: Face-based Speaker Generation by Face-voice Association Learning

  • 利用CLIP模型在噪声数据中学习人脸与语音的语义关联。
  • 在Voxceleb测试集上达到89.63%的跨模态验证准确率。
  • 适合做个性化语音合成、跨模态生成任务的研究者使用。

本文研究基于人脸的语音合成任务,即合成语音需与参考人脸图像在感知上匹配。由于缺乏高质量音视频语料,以往方法或合成质量低,或因知识迁移导致领域不匹配。本文提出Vclip方法,利用CLIP编码器在嘈杂音视频数据中高效学习人脸与语音的关联,在Voxceleb测试集上取得89.63%的跨模态验证AUC得分。该方法采用基于检索的策略,结合基于GMM的说话人生成模块,为参考图像生成潜在目标说话人。实验表明,Vclip系统配合检索步骤能有效弥合人脸与语音特征差距,且从下游TTS中提取反馈信息有助于生成更贴近参考人脸的语音。演示地址:sos1sos2sixteen.github.io/vclip。

原文摘要 · Abstract (English)

This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a reference face image. Due to the lack of TTS-quality audio-visual corpora, previous approaches suffer from either low synthesis quality or domain mismatch induced by a knowledge transfer scheme. This paper proposes a new approach called Vclip that utilizes the facial-semantic knowledge of the CLIP encoder on noisy audio-visual data to learn the association between face and voice efficiently, achieving 89.63% cross-modal verification AUC score on Voxceleb testset. The proposed method then uses a retrieval-based strategy, combined with GMM-based speaker generation module for a downstream TTS system, to produce probable target speakers given reference images. Experimental results demonstrate that the proposed Vclip system in conjunction with the retrieval step can bridge the gap between face and voice features for face-based speech synthesis. And using the feedback information distilled from downstream TTS helps to synthesize voices that match closely with reference faces. Demos available at sos1sos2sixteen.github.io/vclip.

语音合成跨模态人脸匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。