仅凭一张人脸图像,就能生成匹配的语音,无需参考音频。
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
- 用轻量级适配器将人脸特征映射到语音风格空间。
- 合成语音自然度达UTMOS 3.7-4.0,接近真实语音水平。
- 跨语言通用,英文训练可直接生成流畅西语语音。
零样本文本到语音(TTS)依赖短音频提示克隆声音,但在仅能获取视觉信息(如历史人物或游戏角色)时面临障碍。本文提出一种面部到语音(F2S)框架,仅通过静态人脸图像预测合理语音。采用轻量级面部适配器,并软调优人脸编码器上层模块,将人脸识别特征与冻结的StyleTTS 2模型风格空间对齐,训练期间保持模型冻结。在大型音视频数据集LRS3的未见身份上评估,合成语音高度自然(UTMOS 3.7–4.0,与真实语音的3.61相当或更优),人脸-语音检索效果显著高于随机水平,且语音与目标说话人一致。无需重新训练,英文训练的适配器亦能生成流利西班牙语语音,表明该人脸到风格映射具有较强语言无关性。
原文摘要 · Abstract (English)
Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this work, we propose a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image. A lightweight Face Adapter, together with soft-tuning of the face encoder's upper blocks, aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training. We evaluate on held-out identities from LRS3, a large-scale audiovisual corpus of English TED-talk videos. The synthesized speech is highly natural (UTMOS 3.7-4.0, matching or exceeding the 3.61 of ground truth), face-to-voice retrieval is consistently above chance, and the generated voice is consistent with the target speaker. Without any retraining, an English-trained adapter also produces fluent Spanish speech, indicating that the face-to-style mapping is largely language-agnostic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。