arXiv:2501.10045cs.SDeess.AS2025-01中稿 · ICASSP 2025被引 8

用统一模型实现高质量语音超分辨率,提升高音频细节还原能力。

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution

  • 采用统一的变换器-卷积生成器,端到端训练提升音质一致性。
  • 在4kHz~32kHz输入下,可将语音升采至48kHz,显著优于现有方法。
  • 多尺度时频判别器与损失函数设计,增强高频信息还原能力。

生成对抗网络(GAN)在基于中间表示(如梅尔谱图)的语音超分辨率(SR)中取得进展。然而,现有方法通常依赖独立训练并拼接的网络,导致表征不一致,尤其在域外场景下语音质量较差。本文提出HiFi-SR,一种统一的生成式对抗网络,实现高保真语音超分辨率。模型采用统一的变换器-卷积生成器,无缝完成潜在表征预测与时域波形转换:变换器作为强编码器,将低分辨率梅尔谱图映射至潜在空间;卷积网络则将该表示上采样为高分辨率波形。为提升高频保真度,引入多带、多尺度时频判别器,并在对抗训练中加入多尺度梅尔重建损失。HiFi-SR具备通用性,可将4kHz至32kHz的任意输入语音升采至48kHz。实验表明,无论在域内或域外场景,其在客观指标与ABX偏好测试中均显著超越现有方法(https://github.com/modelscope/ClearerVoice-Studio)。

原文摘要 · Abstract (English)

The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent representations and poor speech quality, especially in out-of-domain scenarios. In this work, we propose HiFi-SR, a unified network that leverages end-to-end adversarial training to achieve high-fidelity speech super-resolution. Our model features a unified transformer-convolutional generator designed to seamlessly handle both the prediction of latent representations and their conversion into time-domain waveforms. The transformer network serves as a powerful encoder, converting low-resolution mel-spectrograms into latent space representations, while the convolutional network upscales these representations into high-resolution waveforms. To enhance high-frequency fidelity, we incorporate a multi-band, multi-scale time-frequency discriminator, along with a multi-scale mel-reconstruction loss in the adversarial training process. HiFi-SR is versatile, capable of upscaling any input speech signal between 4 kHz and 32 kHz to a 48 kHz sampling rate. Experimental results demonstrate that HiFi-SR significantly outperforms existing speech SR methods across both objective metrics and ABX preference tests, for both in-domain and out-of-domain scenarios (https://github.com/modelscope/ClearerVoice-Studio).

语音超分辨率生成对抗网络高保真音频统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。