通过分段建模面部细粒度特征,让人脸生成更贴合本人声音。
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
- 将人脸拆成区块逐步融合,保留性别、种族等细微特征
- 多任务学习跨视觉与语音域提升声线一致性,准确率提升12.7%
- 多视角训练增强鲁棒性,适合语音重建与残障人士沟通辅助
中风等创伤后,患者可能失去说话能力。虽可使用文本转语音(TTS)交流,但无法保留个人音色。面部转语音(FTV)从人脸生成对应声音,是更优方案。现有方法依赖预训练视觉编码器并微调对齐语音嵌入,却丢失了性别、种族等与声线相关的细粒度信息。且流程为多阶段,需分别训练多个组件,效率低下。本文提出渐进式面部细粒度聚合方法,将人脸分解为非重叠区域,逐步构建多层次表征,并在视觉与声学域同时进行性别、种族等说话人属性的多任务学习。此外,采用多视角训练策略,用不同角度和光照条件下的人脸图像与相同语音配对,提升对齐鲁棒性。大量主观与客观评估显示,该方法显著提升面部-语音一致性与合成稳定性,声线匹配准确率提高12.7%。
原文摘要 · Abstract (English)
For individuals who have experienced traumatic events such as strokes, speech may no longer be a viable means of communication. While text-to-speech (TTS) can be used as a communication aid since it generates synthetic speech, it fails to preserve the user's own voice. As such, face-to-voice (FTV) synthesis, which derives corresponding voices from facial images, provides a promising alternative. However, existing methods rely on pre-trained visual encoders, and finetune them to align with speech embeddings, which strips fine-grained information from facial inputs such as gender or ethnicity, despite their known correlation with vocal traits. Moreover, these pipelines are multi-stage, which requires separate training of multiple components, thus leading to training inefficiency. To address these limitations, we utilize fine-grained facial attribute modeling by decomposing facial images into non-overlapping segments and progressively integrating them into a multi-granular representation. This representation is further refined through multi-task learning of speaker attributes such as gender and ethnicity at both the visual and acoustic domains. Moreover, to improve alignment robustness, we adopt a multi-view training strategy by pairing various visual perspectives of a speaker in terms of different angles and lighting conditions, with identical speech recordings. Extensive subjective and objective evaluations confirm that our approach substantially enhances face-voice congruence and synthesis stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。