arXiv:2501.03181cs.SDcs.AI2025-01AAAI被引 4

根据不同风格的人像生成匹配性格的高质量语音

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles

  • 从多样化风格人像中提取身份与情绪特征用于语音合成
  • 有效抑制背景、服饰等无关信息干扰,提升语音一致性
  • 适用于虚拟角色、动画配音等需要个性语音的场景

人类可通过外貌感知说话者的特征(如身份、性别、个性和情绪),这些特征通常与语音风格一致。当前视觉驱动的文本转语音研究多基于真实人脸,限制了在多样角色和图像风格场景中的应用。为此,我们提出FaceSpeak方法,能从多种图像风格中提取显著的身份特征与情感表征,并有效消除背景、服装、发色等无关信息,使合成语音更贴近角色人格。为应对多模态语音数据稀缺问题,我们构建了精心标注的表达性多模态语音数据集Expressive Multi-Modal TTS。实验表明,FaceSpeak可生成与人像风格高度对齐、自然度与质量均达满意的语音。

原文摘要 · Abstract (English)

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their investigations on real-person faces, thereby restricting effective speech synthesis from applying to vast potential usage scenarios with diverse characters and image styles. To solve this issue, we introduce a novel FaceSpeak approach. It extracts salient identity characteristics and emotional representations from a wide variety of image styles. Meanwhile, it mitigates the extraneous information (e.g., background, clothing, and hair color, etc.), resulting in synthesized speech closely aligned with a character's persona. Furthermore, to overcome the scarcity of multi-modal TTS data, we have devised an innovative dataset, namely Expressive Multi-Modal TTS, which is diligently curated and annotated to facilitate research in this domain. The experimental results demonstrate our proposed FaceSpeak can generate portrait-aligned voice with satisfactory naturalness and quality.

语音合成人像驱动多模态角色语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。