用语音数据提升跨说话人发音一致性,让不同人发同一词时动作更相似。
Training Articulatory Inversion Models for Interspeaker Consistency
- 基于声学-发音逆映射,用语音数据训练模型预测发音动作
- 单/多说话人训练的模型在英俄语中均实现更高跨说话人一致性
- 新方法仅用语音数据就可改善跨人发音动作的一致性
声学到发音逆映射(AAI)旨在建模从语音到发音运动的逆向映射。仅从语音精确预测发音可能不可行,因为说话人可选择不同发音方式而无需参考其声道结构。然而,一旦选定发音方式,其发音变化极小。近期研究将自监督学习(SSL)模型适配于单说话人数据集,声称这些模型能提供通用的发音模板。本文通过一种新型评估方法,利用最小对集提取发音目标,检验了在单说话人和多说话人数据上训练的SSL模型在英语和俄语中是否具备跨说话人一致性。我们还提出一种仅使用语音数据的训练方法,可显著提升跨说话人发音一致性。
原文摘要 · Abstract (English)
Acoustic-to-Articulatory Inversion (AAI) attempts to model the inverse mapping from speech to articulation. Exact articulatory prediction from speech alone may be impossible, as speakers can choose different forms of articulation seemingly without reference to their vocal tract structure. However, once a speaker has selected an articulatory form, their productions vary minimally. Recent works in AAI have proposed adapting Self-Supervised Learning (SSL) models to single-speaker datasets, claiming that these single-speaker models provide a universal articulatory template. In this paper, we investigate whether SSL-adapted models trained on single and multi-speaker data produce articulatory targets which are consistent across speaker identities for English and Russian. We do this through the use of a novel evaluation method which extracts articulatory targets using minimal pair sets. We also present a training method which can improve interspeaker consistency using only speech data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。