用语言学方法客观评估语音合成的语调,发现传统听感测试看不到的问题。
Toward Objective and Interpretable Prosody Evaluation in Text-to-Speech: A Linguistically Motivated Approach
- 构建双层框架,结合语言学规则与声学特征评估语调
- 与人工评分高度相关,能识别模型特有的语调缺陷
- 适合语音合成研发者诊断和优化语调自然度
语调对语音技术至关重要,影响理解、自然度和表现力。然而当前文本转语音(TTS)系统仍难以准确捕捉类人语调变化,部分原因在于现有语调评估方法有限。传统指标如平均意见分(MOS)耗时耗力、结果不一致,且无法解释为何听起来不自然。本研究提出一种基于语言学的半自动语调评估框架,采用双层架构模拟人类语调组织方式。该方法通过量化语言学标准,在多个声学维度上对比合成语音与真人语音数据库。结合离散与连续语调度量,提供事件位置与语音实现的客观可解释指标,并考虑说话人及语调线索的自然变异。实验显示其结果与感知MOS评分高度相关,同时揭示了传统听感测试无法捕捉的模型特定弱点。该方法为诊断、基准测试和提升下一代TTS系统语调自然度提供了原则性路径。
原文摘要 · Abstract (English)
Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing evaluation methods for prosody remain limited. Traditional metrics like Mean Opinion Score (MOS) are resource-intensive, inconsistent, and offer little insight into why a system sounds unnatural. This study introduces a linguistically informed, semi-automatic framework for evaluating TTS prosody through a two-tier architecture that mirrors human prosodic organization. The method uses quantitative linguistic criteria to evaluate synthesized speech against human speech corpora across multiple acoustic dimensions. By integrating discrete and continuous prosodic measures, it provides objective and interpretable metrics of both event placement and cue realization, while accounting for the natural variability observed across speakers and prosodic cues. Results show strong correlations with perceptual MOS ratings while revealing model-specific weaknesses that traditional perceptual tests alone cannot capture. This approach provides a principled path toward diagnosing, benchmarking, and ultimately improving the prosodic naturalness of next-generation TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。