提出可动态调节频谱分辨率的通用谐波判别器,提升歌声合成质量。
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
- 用可学习的三角带通滤波组实现动态频率分辨率
- 在语音和歌声数据集上主观与客观指标均显著提升
- 特别适合对低频谐波细节敏感的歌声合成任务
随着基于GAN的声码器的发展,判别器作为关键组件受到广泛关注。本文聚焦于时频域判别器的改进,针对传统短时傅里叶变换(STFT)谱图在所有频段具有固定频率分辨率的问题,导致在歌唱语音上表现不佳。为此,我们提出一种通用谐波判别器,支持动态频率分辨率建模与谐波跟踪。具体地,设计了可学习的三角形带通滤波组,使每个频点具备灵活带宽;同时引入半谐波模块,以捕捉低频段精细的谐波关系。在语音和歌唱数据集上的实验表明,所提判别器在主观和客观评价指标上均有效提升。
原文摘要 · Abstract (English)
With the emergence of GAN-based vocoders, the discriminator, as a crucial component, has been developed recently. In our work, we focus on improving the time-frequency based discriminator. Particularly, Short-Time Fourier Transform (STFT) representation is usually used as input of time-frequency based discriminator. However, the STFT spectrogram has the same frequency resolution at different frequency bins, which results in an inferior performance, especially for singing voices. Motivated by this, we propose a universal harmonic discriminator for dynamic frequency resolution modeling and harmonic tracking. Specifically, we design a harmonic filter with learnable triangular band-pass filter banks, where each frequency bin has a flexible bandwidth. Additionally, we add a half-harmonic to capture fine-grained harmonic relationships at low-frequency band. Experiments on speech and singing datasets validate the effectiveness of the proposed discriminator on both subjective and objective metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。