测试神经TTS模型对辅音引发的声调微扰建模能力
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
- 设计分段级韵律探针,评估合成语音中辅音引起的声调变化
- 高频词声调还原准确,低频词表现差,显示模型依赖记忆而非泛化
- 为语音合成真实性评估提供语言学可解释的诊断工具
本研究提出一种分段级韵律探针框架,用于评估神经文本转语音(TTS)模型对辅音引发基频(F0)微扰的建模能力,这是一种反映局部发音机制的细微分段韵律效应。我们在相同语料库(LJ Speech)上训练的Tacotron 2与FastSpeech 2模型基础上,对数千个词汇的合成与自然语音实现进行分层对比,按词频划分。进一步通过涵盖多个先进TTS系统的规模化评估验证。结果表明:高频词的声调微扰还原准确,但对低频词泛化能力差,暗示所考察的TTS架构更依赖词汇级记忆而非抽象的分段韵律编码。该发现揭示了当前TTS系统在泛化韵律细节方面的局限性。提出的探针框架具有语言学启发性,可为未来语音合成评估提供参考,并对合成语音的可解释性与真实感评估具有意义。
原文摘要 · Abstract (English)
This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical frequency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization than on abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems' ability to generalize prosodic detail beyond seen data. The proposed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。