arXiv:2508.01796cs.SDeess.AS2025-08被引 2

通过显式频带扩展提升歌声合成音质真实感

Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder

  • 用基于DiT的去噪扩散模型显式估计线性频谱
  • 新声码器处理更高频点,显著减少高频失真
  • 适合追求高保真歌声合成的研究与开发者

本文针对声码器生成歌声音频时高频频谱成分与真实录音差异明显的问题,提出一种新方法。该方法包含两项创新:一是采用基于DiT架构的去噪扩散模型,显式估计线性频谱;二是设计专用声码器Vocos,可高效处理更高频率分辨率的线性频谱。联合优化后,生成音频的频谱质量极高,人类听者与机器分类器均难以区分其与真实录音的差异。客观与主观评估表明,该方法在保持高音频质量的同时显著提升了真实性。本工作为突破现有声码技术局限,特别是在对抗虚假频谱检测方面提供了重要进展。

原文摘要 · Abstract (English)

This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life recordings, particularly in high-frequency spectrogram components. Our proposed approach combines two innovations: an explicit linear spectrogram estimation step using denoising diffusion process with DiT-based neural network architecture optimized for time-frequency data, and a redesigned vocoder based on Vocos specialized in handling large linear spectrograms with increased frequency bins. This integrated method can produce audio with high-fidelity spectrograms that are challenging for both human listeners and machine classifiers to differentiate from authentic recordings. Objective and subjective evaluations demonstrate that our streamlined approach maintains high audio quality while achieving this realism. This work presents a substantial advancement in overcoming the limitations of current vocoding techniques, particularly in the context of adversarial attacks on fake spectrogram detection.

歌声合成声码器频谱增强扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。