用人脸生成语音,还能分离说话风格与内容,零样本即用。
Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
- 通过面部特征提取说话人专属标识,提升音色匹配精度。
- 音频内容与身份解耦后,语音转换更自然清晰。
- 支持文本或语音输入,可调情感和语速,适合交互应用。
基于人脸的语音转换(FVC)是一项新任务,利用人脸图像生成目标说话人的语音风格。以往方法存在两大问题:一是难以获得与说话人语音身份信息对齐的人脸特征;二是无法有效解耦音频中的内容与说话人身份信息。为此,本文提出一种新型FVC方法——身份解耦式人脸语音转换(ID-FaceVC),以解决上述问题。具体而言,设计了面向身份感知的查询对比学习(IAQ-CL)模块,用于提取具有说话人特异性的面部特征;并引入基于互信息的双重解耦(MIDD)模块,从音频中净化内容特征,确保语音转换的清晰度与高质量。此外,不同于以往方法,本模型可接受音频或文本输入,支持可控语音生成,具备可调节的情感语气与语速。大量实验表明,ID-FaceVC在多项指标上达到当前最优性能,定性分析与用户研究均验证其在自然度、相似性和多样性方面的有效性。项目网站提供音频样本与代码:https://id-facevc.github.io。
原文摘要 · Abstract (English)
Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are well-aligned with the speaker's voice identity information, and (2) inadequacy in decoupling content and speaker identity information from the audio input. To address these issues, we present a novel FVC method, Identity-Disentanglement Face-based Voice Conversion (ID-FaceVC), which overcomes the above two limitations. More precisely, we propose an Identity-Aware Query-based Contrastive Learning (IAQ-CL) module to extract speaker-specific facial features, and a Mutual Information-based Dual Decoupling (MIDD) module to purify content features from audio, ensuring clear and high-quality voice conversion. Besides, unlike prior works, our method can accept either audio or text inputs, offering controllable speech generation with adjustable emotional tone and speed. Extensive experiments demonstrate that ID-FaceVC achieves state-of-the-art performance across various metrics, with qualitative and user study results confirming its effectiveness in naturalness, similarity, and diversity. Project website with audio samples and code can be found at https://id-facevc.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。