arXiv:2609.03357cs.SD2026-09

用双谱图结构提升低质量音乐录音的清晰度

Neural Music Enhancement with Dual Time-Frequency Spectral Representations for Prediction and Discrimination

  • 生成器用STFT谱图重建波形,判别器用CQT谱图识别音高和谐性
  • 在客观与主观测试中均优于基线模型,尤其对混响噪声抑制效果显著
  • 适合音频修复、音乐制作及智能听觉系统开发者参考

非专业音乐录制常因背景噪音和混响导致音质下降,限制其再利用。本文提出基于双时频谱表示的音乐增强模型DSME。在生成对抗框架下,生成器采用短时傅里叶变换(STFT)谱图进行波形重建,利用其固定窗口、可逆性和可预测性,从退化输入估计干净的幅值-相位谱并逆变换回波形;判别器则使用常量-Q变换(CQT)谱图,借助其对数频率、变窗结构与音乐八度对齐的特点,设计分八度的CQT判别器。同时引入音高谱损失以强化音高与谐波一致性。实验表明,DSME在客观指标与主观评测中均优于现有基线,验证了双谱图方法的有效性。

原文摘要 · Abstract (English)

Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT's fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT's log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.

音乐增强生成对抗频谱表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。