arXiv:2605.00861eess.AScs.AI2026-05

用声音映射评估语音合成质量,发现VITS音域最广、自然度与参数相关。

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment

  • 通过频谱平衡、峰度等指标构建声音映射框架
  • CPPs在7-8dB时更自然,超10dB则显机械感
  • 适合关注语音动态与表现力的语音系统研究者

本研究探讨声音映射作为文本转语音(TTS)合成质量的评估框架。分析了六种有影响力的TTS模型:Merlin、Tacotron 2、Transformer TTS、FastSpeech 2、Glow-TTS和VITS。采用峰值因子、频谱平衡和倒谱峰度突出(CPPs)作为评估指标。结果表明,音域是模型能力的主要指标,其中VITS在测试模型中音域最广。尽管音域有限,Glow-TTS在轻声发音方面表现更优,体现更高的频谱平衡。当CPPs值在7-8 dB之间时,语音更接近自然人声;而超过10 dB时,语音趋于机械化。这些发现强调了声音映射在评估语音努力程度及捕捉TTS系统对语音动态与表现力处理能力方面的重要性。

原文摘要 · Abstract (English)

This study investigates voice mapping as an evaluation framework for text-to-speech (TTS) synthesis quality. The study analyzes six TTS models, including historical and recent ones. The metrics are crest factor, spectrum balance, and cepstral peak prominence (CPPs). We investigated 6 influential TTS models: Merlin, Tacotron 2, Transformer TTS, FastSpeech 2, Glow-TTS, and VITS. The results demonstrate that voice range serves as a primary indicator of model capability, with VITS showing the largest range among tested models. Glow-TTS exhibited superior performance in soft phonation, indicated by higher spectrum balance, despite limited voice range. The results showed that the CPPs values between 7-8 dB indicate natural voice quality, while with CPPs exceeding 10 dB, the speech tends to sound robotic. These findings underscore the need for voice mapping to evaluate vocal effort, and capture how TTS systems handle voice dynamic and expressiveness.

语音合成质量评估声音映射自然度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。