arXiv:2607.01965cs.CLcs.ET2026-07中稿 · Interspeech 2026

用语音分类器检测多语言语音合成是否保留音位差异,发现合成音有偏移。

Towards a Phonology-Informed Evaluation of Multilingual TTS

  • 用人类语音训练分类器,检测合成语音的音位特征是否准确
  • 三分之一带[+ATR]标记的元音被错误合成成[-ATR],人类无此偏差
  • 适合关注语音合成音位保真度的研究者或开发者

神经语音合成系统在多种语言中可实现自然发音,但自然性并不保证保留区分词义与语法形式的声音对比。标准指标如MOS无法检验这一点。我们提出一种基于分类器的评估框架,利用人类语音作为基准,审计合成语音是否符合特定语言的音系模式。以梅泰语的前舌根(ATR)元音和谐为例,测试Meta的MMS TTS系统发现,基于人类语音训练的分类器在合成语音上仅略有性能下降。忠实度审计揭示:尽管底层标注为[+ATR],仍有1/3的中元音被错误合成成[-ATR],而人类语音中不存在此偏差。在词级层面,预测的ATR标签比转写标签更准确地反映和谐规律,表明合成语音与预期音位存在差距。该框架提供任务特定诊断,可推广至其他具有可测量声学线索的音位对比。

原文摘要 · Abstract (English)

Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this. We propose a classifier-based framework that audits TTS output against language-specific phonological patterns using human speech as a benchmark. Testing Assamese advanced tongue root (ATR) vowel harmony with Meta's MMS TTS, we show that a classifier trained on human speech transfers to synthesized speech with minimal loss. The faithfulness audit reveals that [+ATR] mid vowels are realized as [-ATR] in 1/3 tokens despite an underlying [+ATR] specification, a bias absent in human speech. At the word level, predicted ATR labels classify harmony more accurately than transcription labels, indicating a gap between intended and produced phonology. The framework offers task-specific diagnostics and generalizes to other phonological contrasts with measurable acoustic cues.

语音合成音位保真多语言TTS评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。