arXiv:2512.14653cs.SD2025-12中稿 · ASRU 2025被引 1

通过不确定性优化提升歌声合成在数据稀缺下的鲁棒性

Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty

  • 用可微数据增强增加先验不确定性,提升模型泛化能力
  • 引入帧级后验不确定性预测,聚焦低置信度片段训练
  • 在中日语种的公开数据集上均显著提升长尾场景性能

近年来,歌声合成(SVS)取得显著进展。然而,与语音和通用音频数据相比,公开的歌唱数据集仍十分有限。实践中,这种数据稀缺常导致长尾场景下性能下降,如音高分布不均或罕见演唱风格。为缓解此问题,我们提出基于不确定性的优化方法,改进端到端SVS模型的训练过程。首先,在对抗训练中引入可微数据增强,以样本级方式增加先验不确定性;其次,加入帧级不确定性预测模块,估计后验不确定性,使模型能将更多学习容量分配给低置信度段。在中文与日文的Opencpop和Ofuton-P数据集上的实验表明,该方法在多个维度上均有效提升性能。

原文摘要 · Abstract (English)

Singing voice synthesis (SVS) has seen remarkable advancements in recent years. However, compared to speech and general audio data, publicly available singing datasets remain limited. In practice, this data scarcity often leads to performance degradation in long-tail scenarios, such as imbalanced pitch distributions or rare singing styles. To mitigate these challenges, we propose uncertainty-based optimization to improve the training process of end-to-end SVS models. First, we introduce differentiable data augmentation in the adversarial training, which operates in a sample-wise manner to increase the prior uncertainty. Second, we incorporate a frame-level uncertainty prediction module that estimates the posterior uncertainty, enabling the model to allocate more learning capacity to low-confidence segments. Empirical results on the Opencpop and Ofuton-P, across Chinese and Japanese, demonstrate that our approach improves performance in various perspectives.

歌声合成不确定性建模数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。