用物理约束神经算子快速模拟发音过程,无需训练数据。
Physics-Informed Neural Operator for Speech Production Analysis
- 直接学习声带波动方程,不依赖标注数据
- 音高与声波误差分别仅0.8%和3.2%
- 适合语音分析与逆问题求解场景
物理信息神经算子(PINOs)作为快速数值模拟器,近年来在求解逆问题方面展现出潜力。本文首次提出基于PINOs的发音生成分析方法,模型直接学习一维波动方程,无需预训练监督数据。以声道形状为输入,对比了五种静态元音的基频、声门体积流速及唇端声压预测结果。相比传统龙格-库塔/有限差分方法,该模型在声门体积流速上误差仅为0.8%,语音波形误差为3.2%,实现高效的GPU并行仿真,无需迭代计算。结果表明,PINO是语音快速分析的有力候选方法。
原文摘要 · Abstract (English)
Physics-informed neural operators (PINOs) have recently gained attention as fast numerical simulators with potential for solving inverse problems. This study proposes the first PINO-based method for speech production analysis. The model learns the governing one-dimensional wave equations directly without requiring pre-computed supervised training data. Using vocal tract shape data as input features, we compare the proposed model's predicted f0, glottal volume velocity and sound pressure at the lip for five static vowels to a conventional Runge Kutta/Finite difference approach. With errors of 0.8% for glottal volume flow and 3.2% for speech waveforms, the proposed model enables efficient GPU-parallelized simulation without iterative calculations. We conclude that PINO is a promising approach for fast analysis of speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。