arXiv:2601.00217cs.SDcs.AI2026-01

通过流匹配修复生成时的潜在表示偏差,让歌声更自然

Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching

  • 用流匹配在潜空间中优化生成时的潜在向量
  • 在韩语和汉语数据集上提升音质与感知评分
  • 不改变原有模型,适合想提升歌声表现力的研究者

歌唱语音合成(SVS)旨在从符号音乐谱生成自然且富有表现力的歌唱波形。然而,在条件变分自编码器(cVAE)基的SVS中,解码器使用从目标歌唱信号推断出的潜在表示进行训练,而推理时仅依赖于条件输入预测的潜在表示,这种差异会削弱合成输出中的精细表现性声学细节。为缓解此问题,我们提出FM-Singer——一种基于流匹配的潜空间精炼框架。该方法不重设计声学解码器,而是学习一个连续向量场,通过常微分方程(ODE)积分将推理时的潜在样本逐步引导至后验类潜表示。由于精炼在潜空间进行,该方法轻量且兼容强的并行合成主干网络。在韩语和汉语歌唱数据集上的实验表明,所提潜空间精炼提升了客观指标与主观质量,同时保持了高效的合成性能。结果表明,减少训练-推理间的潜空间失配是提升表现性歌唱语音合成的有效方向。代码、预训练检查点和音频演示见https://github.com/alsgur9368/FM-Singer。

原文摘要 · Abstract (English)

Singing voice synthesis (SVS) aims to generate natural and expressive singing waveforms from symbolic musical scores. In cVAE-based SVS, however, a mismatch arises because the decoder is trained with latent representations inferred from target singing signals, while inference relies on latent representations predicted only from conditioning inputs. This discrepancy can weaken fine expressive acoustic details in the synthesized output. To mitigate this issue, we propose FM-Singer, a flow-matching-based latent refinement framework for cVAE-based singing voice synthesis. Rather than redesigning the acoustic decoder, the proposed method learns a continuous vector field that transports inference-time latent samples toward posterior-like latent representations through ODE-based integration before waveform generation. Because the refinement is performed in latent space, the method remains lightweight and compatible with a strong parallel synthesis backbone. Experimental results on Korean and Chinese singing datasets show that the proposed latent refinement improves objective metrics and perceptual quality while maintaining practical synthesis efficiency. These results suggest that reducing training-inference latent mismatch is a useful direction for improving expressive singing voice synthesis. Code, pre-trained checkpoints, and audio demos are available at https://github.com/alsgur9368/FM-Singer.

语音合成潜空间流匹配表达力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。