让文字生成的说话人脸与声音自动匹配,还能灵活控制音调语调。
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
- 通过语音-人脸对齐机制,确保生成声音与面部特征一致。
- 在单块40GB显卡上实现视觉与音频双领先效果。
- 轻量化设计适合快速部署,适合需要语音可控的应用场景。
近期基于语音驱动的说话人脸生成研究取得了良好效果,但其依赖固定语音输入限制了进一步应用(如人脸与声音不匹配)。为此,我们拓展任务至更具挑战性的设定:给定一张人脸图像和待说文本,同时生成对应的说话人脸动画及其语音。为此提出新框架 Face2VoiceSync,包含四项创新:1)语音-人脸对齐,确保生成语音与面部外观一致;2)多样性与可操控性,支持对副语言特征空间的语音控制;3)高效训练,采用轻量级VAE连接视觉与音频大模型,参数量显著低于现有方法;4)新评估指标,公平衡量生成多样性与身份一致性。实验表明,Face2VoiceSync 在单块40GB GPU上实现了视觉与音频的当前最优性能。
原文摘要 · Abstract (English)
Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more challenging setting: given a face image and text to speak, generating both talking face animation and its corresponding speeches. Accordingly, we propose a novel framework, Face2VoiceSync, with several novel contributions: 1) Voice-Face Alignment, ensuring generated voices match facial appearance; 2) Diversity \& Manipulation, enabling generated voice control over paralinguistic features space; 3) Efficient Training, using a lightweight VAE to bridge visual and audio large-pretrained models, with significantly fewer trainable parameters than existing methods; 4) New Evaluation Metric, fairly assessing the diversity and identity consistency. Experiments show Face2VoiceSync achieves both visual and audio state-of-the-art performances on a single 40GB GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。