仅凭人脸视频实现无声说话人语音转换,让静默视频开口说目标人声音。
MuteSwap: Visual-informed Silent Video Identity Conversion
- 用视觉信息替代音频,通过对比学习对齐跨模态说话人身份
- 在噪声环境下仍能生成可懂语音,音色转换准确率显著提升
- 适合无音频或音频质量差的场景,如监控视频修复、隐私保护
传统语音转换依赖源和目标说话人的清晰音频,但在无声视频或嘈杂环境中无法使用。本文聚焦无声面部语音转换(SFVC)任务:仅基于目标说话人图像和源说话人静默视频中的口型运动,生成符合目标说话人身份但保留源内容的语音。该任务需仅凭视觉线索生成可懂语音并完成身份转换,极具挑战性。为此,我们提出MuteSwap框架,采用对比学习对齐跨模态身份,并最小化互信息以分离共享视觉特征。实验表明,MuteSwap在语音合成与身份转换上表现优异,尤其在噪声环境下,依赖音频的方法失效时仍能输出可懂结果,验证了训练方法的有效性与SFVC的可行性。
原文摘要 · Abstract (English)
Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。