通过语音感知风格提取与方向调整,提升语音合成表现力。
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- 聚焦语音段落中的风格相关区域,增强表达性。
- 风格方向调整使合成语音质量显著提升。
- 适合追求自然情感语音的语音合成研究者。
近期的富有表现力的文本转语音(TTS)技术多基于从参考语音中提取的风格嵌入。然而,生成高质量的有表现力语音仍具挑战。本文提出Spotlight-TTS,通过语音感知风格提取与风格方向调整,专注于语音中的风格信息。该方法聚焦于与风格高度相关的有声区域,同时保持不同语音区域间的连续性,以提升表达力;并通过调整提取风格的方向,实现更优的集成效果。实验结果表明,Spotlight-TTS在表现力、整体语音质量及风格迁移能力方面均优于基线模型。音频样本已公开。
原文摘要 · Abstract (English)
Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose Spotlight-TTS, which exclusively emphasizes style via voiced-aware style extraction and style direction adjustment. Voiced-aware style extraction focuses on voiced regions highly related to style while maintaining continuity across different speech regions to improve expressiveness. We adjust the direction of the extracted style for optimal integration into the TTS model, which improves speech quality. Experimental results demonstrate that Spotlight-TTS achieves superior performance compared to baseline models in terms of expressiveness, overall speech quality, and style transfer capability. Our audio samples are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。