arXiv:2504.18099cs.SDcs.CL2025-04

用固定权重的双向LSTM-CNN预测语音中的舌位和唇部动作

Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture

  • 采用固定权重的1D CNN后处理双向LSTM,提升序列建模效率
  • 在跨说话人、跨语料场景下仍保持较高预测准确率
  • 适合语音合成与发音机制研究,尤其关注声道动态

语音产生是一个复杂的序列过程,涉及多种发音特征的协调。其中舌部作为高度灵活的主动发音器官,负责调节气流以生成清晰、明确的语音。本文提出一种新方法,通过堆叠双向长短期记忆网络(BiLSTM)结合一维卷积神经网络(CNN)进行后处理,利用固定权重初始化,从给定语音声学信号中预测舌部与唇部的发音特征。模型在两个同步采集语音与电磁发音图(EMA)的数据集上训练,涵盖不同地理来源、语言特征、音素多样性及录音设备差异。在说话人相关(SD)、说话人无关(SI)、语料相关(CD)和跨语料(CC)模式下评估性能。实验表明,固定权重方法在较少训练轮次下表现优于可调权重初始化。该成果有助于构建鲁棒高效的发音特征预测模型,推动语音生成与发音机制研究进展。

原文摘要 · Abstract (English)

Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech sounds that are intellectual, clear, and distinct. This paper presents a novel approach for predicting tongue and lip articulatory features involved in a given speech acoustics using a stacked Bidirectional Long Short-Term Memory (BiLSTM) architecture, combined with a one-dimensional Convolutional Neural Network (CNN) for post-processing with fixed weights initialization. The proposed network is trained with two datasets consisting of simultaneously recorded speech and Electromagnetic Articulography (EMA) datasets, each introducing variations in terms of geographical origin, linguistic characteristics, phonetic diversity, and recording equipment. The performance of the model is assessed in Speaker Dependent (SD), Speaker Independent (SI), corpus dependent (CD) and cross corpus (CC) modes. Experimental results indicate that the proposed model with fixed weights approach outperformed the adaptive weights initialization with in relatively minimal number of training epochs. These findings contribute to the development of robust and efficient models for articulatory feature prediction, paving the way for advancements in speech production research and applications.

语音生成发音建模深度学习序列预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。