arXiv:2607.01594eess.AS2026-07被引 1

用三类标签预训练,让语音转发音器官动作更高效准确。

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

  • 用音素、发音特征和关键发音体标签联合预训练模型。
  • 低资源下性能提升明显,推理速度更快且不降精度。
  • 适合语音合成、识别等需高效建模的场景。

语音到发音器官映射(AAI)从语音中估计声道发音器官运动,对自动语音识别、语音合成和说话人验证等任务有益。尽管基于深度学习的方法(如CNN、RNN、Transformer)已推动该领域发展,近期研究发现自监督学习(SSL)特征可进一步提升性能,尤其在低资源场景下。然而,SSL特征提取器引入了推理延迟和计算开销。为此,我们提出一种新型预训练方法,利用三种目标表示——音素标签、发音特征标签和关键发音体标签——在推理阶段无需依赖SSL提取器。我们在多种数据条件下对比基线与基于SSL的模型,结果表明该方法在低资源场景下持续提升AAI性能,同时显著降低推理成本,且不损失准确性。

原文摘要 · Abstract (English)

Acoustic-to-Articulatory Inversion (AAI) estimates vocal tract articulator movements from speech, benefiting tasks like ASR, speech synthesis, and speaker verification. While deep learning-based methods (CNNs, RNNs, Transformers) have advanced AAI, recent studies show that Self-Supervised Learning (SSL) features further enhance performance, particularly in low-resource settings. However, SSL feature extractors introduce inference latency and computational overhead. To address this, we propose a novel pretraining method leveraging three target representations-Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels-eliminating the need for an SSL extractor during inference. We evaluate our approach against both baseline and SSL-based models across various data conditions. Results demonstrate that our method consistently improves AAI performance, particularly in low-resource scenarios, while significantly reducing inference costs without sacrificing accuracy.

语音生成低资源预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。