arXiv:2511.14824cs.SDcs.AI2025-11

通过关注语音中的有声区,提升语音合成的表达力与质量。

Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

  • 仅聚焦有声区域提取风格特征,增强表达性
  • 调整风格方向使合成语音更自然,音质更优
  • 适合需要高情感表达的语音合成应用

近期的富有表现力的文本到语音(TTS)技术多依赖从参考语音中提取风格嵌入。然而,生成高质量的有表现力语音仍具挑战。本文提出SpotlightTTS,通过有声感知风格提取与风格方向调整来强化风格建模。有声感知风格提取聚焦于与风格高度相关的有声区域,同时保持不同语音段落间的连续性,以提升表达力;并通过调整提取风格的方向,实现更优地融入TTS模型,显著改善语音质量。实验结果表明,Spotlight-TTS在表达力、整体语音质量及风格迁移能力方面均优于基线模型。

原文摘要 · Abstract (English)

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose SpotlightTTS, which exclusively emphasizes style via voiced-aware style extraction and style direction adjustment. Voiced-aware style extraction focuses on voiced regions highly related to style while maintaining continuity across different speech regions to improve expressiveness. We adjust the direction of the extracted style for optimal integration into the TTS model, which improves speech quality. Experimental results demonstrate that Spotlight-TTS achieves superior performance compared to baseline models in terms of expressiveness, overall speech quality, and style transfer capability.

语音合成风格提取有声感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。