为印地语语音合成设计可解释的发音维度评估基准,精准识别非本土化特征。
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
- 将发音差异分解为6个可量化维度,基于声学嵌入与母语者中心点比对
- 发现泰卢固语和泰米尔语的卷舌音退化率高达40%和68%,远超印地语
- 揭示商业系统在发音和韵律上无一全优,适合多维度评估语音合成质量
标准语音合成评估关注可懂度(WER、CER)与自然度(MOS、UTMOS),但无法量化口音。合成系统在四项指标中表现良好,仍可能在目标语言特有的音位特征上显得非本土化。针对印地语系语言,这些特征包括卷舌发音、送气、元音长短及泰米尔语卷舌近音(字母zha)。我们提出PSP(音素替换谱图),一种面向印地语语音合成的可解释性、逐音位维度口音评估基准。PSP将口音分解为六个互补维度:卷舌退化率(RR)、送气保真度(AF)、元音长短保真度(LF)、泰米尔-zha保真度(ZF)、Frechet音频距离(FAD)和韵律特征偏离度(PSD)。前四个维度通过强制对齐结合Wav2Vec2-XLS-R第9层嵌入的母语者中心声学探针测量;后两个为语料级分布距离。本版本对四个商用与开源系统(ElevenLabs v3、Cartesia Sonic-3、Sarvam Bulbul、Indic Parler-TTS)在印地语、泰卢固语、泰米尔语试点集上进行评估,第五个系统(Praxy Voice)覆盖所有三语,并补充泰卢固语从R5到R6的案例研究。三项发现:(i) 卷舌退化随音位难度递增:印地语<泰卢固语<泰米尔语(约1%、~40%、~68%);(ii) PSP排序与WER排序不一致——商业系统在卷舌或韵律保真度上并非全面领先;(iii) 无单一系统在六维上均最优。我们公开母语参考中心点(每语言500段)、1000段嵌入用于FAD、500段韵律特征矩阵用于PSD、每语言300句黄金语料集、评分代码(MIT许可)、中心点数据(CC-BY许可)。正式的MOS相关性延至v2;v1报告五项内部一致性信号及母语音频合理性检查。
原文摘要 · Abstract (English)
Standard text-to-speech (TTS) evaluation measures intelligibility (WER, CER) and overall naturalness (MOS, UTMOS) but does not quantify accent. A synthesiser may score well on all four yet sound non-native on features that are phonemic in the target language. For Indic languages, these features include retroflex articulation, aspiration, vowel length, and the Tamil retroflex approximant (letter zha). We present PSP, the Phoneme Substitution Profile, an interpretable, per-phonological-dimension accent benchmark for Indic TTS. PSP decomposes accent into six complementary dimensions: retroflex collapse rate (RR), aspiration fidelity (AF), vowel-length fidelity (LF), Tamil-zha fidelity (ZF), Frechet Audio Distance (FAD), and prosodic signature divergence (PSD). The first four are measured via forced alignment plus native-speaker-centroid acoustic probes over Wav2Vec2-XLS-R layer-9 embeddings; the latter two are corpus-level distributional distances. In this v1 we benchmark four commercial and open-source systems (ElevenLabs v3, Cartesia Sonic-3, Sarvam Bulbul, Indic Parler-TTS) on Hindi, Telugu, and Tamil pilot sets, with a fifth system (Praxy Voice) included on all three languages, plus an R5->R6 case study on Telugu. Three findings: (i) retroflex collapse grows monotonically with phonological difficulty Hindi < Telugu < Tamil (~1%, ~40%, ~68%); (ii) PSP ordering diverges from WER ordering -- commercial WER-leaders do not uniformly lead on retroflex or prosodic fidelity; (iii) no single system is Pareto-optimal across all six dimensions. We release native reference centroids (500 clips per language), 1000-clip embeddings for FAD, 500-clip prosodic feature matrices for PSD, 300-utterance golden sets per language, scoring code under MIT, and centroids under CC-BY. Formal MOS-correlation is deferred to v2; v1 reports five internal-consistency signals plus a native-audio sanity check.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。