用物理模型合成歌声,让声音更自然且可解释。
Learning Vocal-Tract Area and Radiation with a Physics-Informed Webster Model
- 用时域韦伯斯特方程建模声道形状和辐射特性。
- 在元音上生成频谱包络,性能媲美紧凑型DDSP基线。
- 适合对语音合成可解释性有要求的研究者。
我们提出一种基于物理的发声后端渲染器,用于歌唱语音合成。给定单通道合成音频和基频轨迹,训练一个时域韦伯斯特模型作为物理信息神经网络,以估计可解释的声道截面积函数和开放端辐射系数。训练过程强制满足偏微分方程与边界一致性;仅使用轻量级DDSP路径稳定学习,推理完全基于物理。在持续元音(/a/、/i/、/u/)上,由独立有限差分时域韦伯斯特求解器生成的参数,其频谱包络表现与紧凑型DDSP基线相当,并在离散化变化、适度声源扰动及约10%基频偏移下保持稳定。生成波形仍略显气声,提示未来需引入周期感知目标与显式声门先验。
原文摘要 · Abstract (English)
We present a physics-informed voiced backend renderer for singing-voice synthesis. Given synthetic single-channel audio and a fund-amental--frequency trajectory, we train a time-domain Webster model as a physics-informed neural network to estimate an interpretable vocal-tract area function and an open-end radiation coefficient. Training enforces partial differential equation and boundary consistency; a lightweight DDSP path is used only to stabilize learning, while inference is purely physics-based. On sustained vowels (/a/, /i/, /u/), parameters rendered by an independent finite-difference time-domain Webster solver reproduce spectral envelopes competitively with a compact DDSP baseline and remain stable under changes in discretization, moderate source variations, and about ten percent pitch shifts. The in-graph waveform remains breathier than the reference, motivating periodicity-aware objectives and explicit glottal priors in future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。