用可微波导模型精准复制图瓦喉音唱法的高音泛音。
Differentiable Articulatory Copy-Synthesis of Biphonic Singing

- 引入次舌第二声源与分段可学习阻尼,提升声道建模精度。
- 在20段音频上降低30%-38%频谱误差,尤在1-3kHz泛音区表现突出。
- 适合语音合成、音乐生成领域研究者,尤其关注泛音控制的场景。
Sygyt是图瓦风格的双声部歌唱,通过持续低音基频的同时,在1-3 kHz区域选择性放大某个高次泛音实现。传统声道参数化方法因难以精确调控窄带共振,难以复现此效果。本文提出一种可微分的Kelly--Lochbaum波导模型,集成次舌第二声源、三次B样条声道参数化及空间可变学习型阻尼,通过梯度下降端到端优化音频信号。在来自两个独立数据集(5位歌手,10个音高)共20段样本上,该模型相较基线人工声道模型降低30%-38%对数谱距离,最大增益集中在泛音区域。倒谱包络分析显示,其能更准确还原sygyt发声特有的合并共振峰结构。此外,该模型优于直接控制每条谐波的DDSP谐波+噪声基线,表明显式声学结构对泛音歌唱复制具有重要归纳偏置作用。
原文摘要 · Abstract (English)
Sygyt is a Tuvan style of biphonic singing in which a low vocal drone is sustained while a high harmonic is selectively amplified in the 1--3\,kHz region. Copy-synthesizing this effect remains challenging for articulatory models, since it requires fine control of narrowly focused resonances that standard low-dimensional tract parameterizations cannot easily reproduce. We address this problem with a differentiable Kelly--Lochbaum waveguide augmented with a sublingual second source, cubic B-spline tract parameterization, and spatially varying learnable damping, optimized end-to-end by gradient descent from audio. On 20 segments from two independent sygyt datasets (5 singers, 10 pitches), the proposed model reduces log-spectral distance by 30--38\% relative to an articulatory baseline, with the largest gains concentrated in the overtone region. Cepstral-envelope analysis further shows more accurate recovery of the merged formant structure characteristic of sygyt production. The model also outperforms a DDSP harmonic-plus-noise baseline with direct per-harmonic spectral control, suggesting that explicit acoustic structure is a useful inductive bias for overtone-singing copy-synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。