arXiv:2508.11273eess.AS2025-08

用球面向量与离散语音标记实现多语言情感语音合成,提升自然度和情感控制力。

EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens

  • 将情绪编码为球面坐标,实现细腻情感控制
  • 在英日语料上显著提升语音清晰度与韵律一致性
  • 适合需要跨语言情感语音生成的研究与应用

本文提出EmoSSLSphere,一种结合球面情绪向量与自监督学习(SSL)提取的离散语音标记的多语言情感文本到语音(TTS)合成框架。通过在连续球面坐标空间中编码情绪,并利用基于SSL的表示进行语义与声学建模,该方法实现了精细的情感控制、有效的跨语言情感迁移以及说话人身份的稳健保持。我们在英语和日语语料上进行评估,结果表明其在语音可懂度、谱保真度、韵律一致性及整体合成质量方面均有显著提升。主观评价进一步证实,相比基线模型,本方法在自然度与情感表现力上更具优势,展现出作为可扩展多语言情感TTS解决方案的巨大潜力。

原文摘要 · Abstract (English)

This paper introduces EmoSSLSphere, a novel framework for multilingual emotional text-to-speech (TTS) synthesis that combines spherical emotion vectors with discrete token features derived from self-supervised learning (SSL). By encoding emotions in a continuous spherical coordinate space and leveraging SSL-based representations for semantic and acoustic modeling, EmoSSLSphere enables fine-grained emotional control, effective cross-lingual emotion transfer, and robust preservation of speaker identity. We evaluate EmoSSLSphere on English and Japanese corpora, demonstrating significant improvements in speech intelligibility, spectral fidelity, prosodic consistency, and overall synthesis quality. Subjective evaluations further confirm that our method outperforms baseline models in terms of naturalness and emotional expressiveness, underscoring its potential as a scalable solution for multilingual emotional TTS.

情感合成多语言TTS球面向量自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。