针对低资源语言设计多维度语音合成评估框架,揭示情感语音最难合成。
Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

- 构建跨语域的主观+客观联合评估体系,涵盖四类语音场景。
- 情感语音合成误差最大(均值MCD 12.03 dB,F0 RMSE 889 cents)。
- 开源完整工具链,支持低资源语言语音合成可复现研究。
近年来神经文本转语音(TTS)系统在多种语言中显著提升了语音自然度与可懂度。然而,针对低资源及未充分代表语言,能够联合评估感知质量、说话人相似性与声学保真度的综合性评估方法仍较匮乏。本文提出一种可复现的多指标基准评估框架,通过领域特定分析系统评估现代TTS系统。该框架融合主观与客观评估协议,在代表性低资源语言上开展全面案例研究,覆盖正式、对话、文学/讲故事、情感四种语音领域。评估了四种先进TTS系统:Indic-Parler-TTS、MMS-TTS、Microsoft Edge TTS和Google Gemini TTS,采用MUSHRA听感测试、ABX辨识测试、基于Resemblyzer的说话人相似性评分,以及基于梅尔倒谱失真(MCD)和基频均方根误差(F0 RMSE)的声学分析,共处理960个音频对。结果表明各语音领域间性能差异显著,情感语音始终是合成难点(均值MCD 12.03 dB;均值F0 RMSE 889 cents),而对话语音整体声学保真度最高。本研究不仅提供实证发现,还公开发布评估脚本、结果表格与可执行的Colab笔记本,支持标准基准测试及未来低资源语言语音合成评估研究。
原文摘要 · Abstract (English)
Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However, comprehensive evaluation methodologies that jointly assess perceptual quality, speaker similarity, and acoustic fidelity across diverse speech domains remain limited, particularly for low-resource and underrepresented languages. This paper presents a reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis. The proposed framework integrates complementary subjective and objective evaluation protocols and is demonstrated through a comprehensive case study on a representative low-resource language spanning four speech domains: Formal, Conversational, Literary/Storytelling, and Emotional. Four state-of-the-art TTS systems -- Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, and Google Gemini TTS -- are evaluated using MUSHRA listening tests, ABX discrimination tests, speaker similarity scoring with Resemblyzer, and acoustic analyses based on mel-cepstral distortion (MCD) and F0 RMSE over 960 audio pairs. Results reveal substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge (mean MCD 12.03 dB; mean F0 RMSE 889 cents), while conversational speech achieves the highest overall acoustic fidelity. Beyond the empirical findings, this work provides a reproducible evaluation framework, publicly releasing evaluation scripts, result tables, and executable Colab notebooks to support standardized benchmarking and future research on TTS evaluation for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。