用音位识别评估发音合成质量,更准确捕捉发音细节。
Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition
- 用发音特征做音位识别,替代传统距离度量。
- 在单人RT-MRI数据上训练模型,识别准确率提升显著。
- 适合语音合成、发音研究方向的学者参考。
机器学习进展与发音数据集的可用性使得声带合成可基于音位序列进行条件控制,这是发音语音合成的核心任务。然而,质量评估仍缺乏明确标准。通常,生成模型的排序因主观性而困难,而发音合成还需具备声道解剖学与声学的专业知识。为此,本文提出以音位识别作为发音合成质量的代理评估指标。假设使用发音特征的音位识别能更好捕捉音位产生的细微差别,如正确的发音部位,而传统点对点距离度量无法做到这一点。我们在单人RT-MRI数据集上提取声学与发音特征,并训练神经网络。随后,通过测试不同合成发音特征下的识别性能进行比较。结果表明,该发音特征集具有丰富的音位信息,有助于探索发音合成的更多维度。
原文摘要 · Abstract (English)
Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality assessment needs a better definition. Generally, ranking generative models is tricky due to subjectivity. However, articulatory synthesis has the additional difficulty of requiring specialized knowledge in vocal tract anatomy and acoustics. To address this problem, this paper proposes to evaluate speech articulation synthesis using phoneme recognition as a proxy. Our hypothesis is that phoneme recognition using articulatory features better captures nuances in phoneme production, such as correct places of articulation, which traditional metrics (e.g., point-wise distance metrics) do not. We train a neural network with acoustic and articulatory features extracted from a single-speaker RT-MRI dataset. Then, we compare the recognition performance when testing the model with different synthetic articulatory features. Our results show that our articulatory feature set is phonetically rich and helps exploring additional dimensions on speech articulation synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。