用6维发音特征建模口腔运动,实现高保真语音合成与逆向重建。
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
- 基于真实MRI数据构建6维发音特征,涵盖软腭、舌根等关键部位。
- 通过声学预测发音特征,相关性达0.87,实现高效逆向建模。
- 仅用低维特征即可合成自然语音,适合语音技术研究者使用。
我们提出ARTI-6,一种基于实时MRI数据的紧凑六维发音编码框架,能够捕捉包括软腭、舌根和喉部在内的关键声道区域。ARTI-6包含三个组成部分:(1) 六维发音特征集,表征声道关键区域;(2) 发音逆向模型,利用语音基础模型从语音声学中预测发音特征,达到0.87的预测相关性;(3) 发音合成模型,可直接从发音特征重建出可理解的语音,表明即使低维表示也能生成自然语音。ARTI-6共同提供了一种可解释、计算高效且生理上合理的框架,推动发音逆向、合成及更广泛语音技术应用的发展。源代码与语音样本已公开。
原文摘要 · Abstract (English)
We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components: (1) a six-dimensional articulatory feature set representing key regions of the vocal tract; (2) an articulatory inversion model, which predicts articulatory features from speech acoustics leveraging speech foundation models, achieving a prediction correlation of 0.87; and (3) an articulatory synthesis model, which reconstructs intelligible speech directly from articulatory features, showing that even a low-dimensional representation can generate natural-sounding speech. Together, ARTI-6 provides an interpretable, computationally efficient, and physiologically grounded framework for advancing articulatory inversion, synthesis, and broader speech technology applications. The source code and speech samples are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。