arXiv:2503.16989cs.SDeess.AS2025-03中稿 · ICME 2025被引 6

用频域表示提升音频压缩质量,兼顾高保真与灵活调节。

STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation

  • 基于短时傅里叶变换构建频域编码器,分离处理幅度与相位分支。
  • 在多个码率下超越波形与频域方法,相位解压更稳定。
  • 调整STFT参数即可改变压缩比,无需改动模型结构。

我们提出STFTCodec,一种基于频谱的新型神经音频编解码器,利用短时傅里叶变换(STFT)实现高效音频压缩。与依赖波形建模、需大模型容量和高内存的方案不同,该方法通过STFT生成紧凑的频谱表示,并引入解缠绕相位导数作为辅助特征。架构采用并行的幅度与相位处理分支,结合先进的特征提取机制。通过放宽严格的相位重建约束,同时保持相位感知处理,实现了更优的主观听感质量。实验表明,STFTCodec在多个码率下均优于波形基与频域基方法,且可通过调整STFT参数灵活调节压缩比,无需修改网络结构。

原文摘要 · Abstract (English)

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory consumption, this method leverages STFT for compact spectral representation and introduces unwrapped phase derivatives as auxiliary features. Our architecture employs parallel magnitude and phase processing branches enhanced by advanced feature extraction mechanisms. By relaxing strict phase reconstruction constraints while maintaining phase-aware processing, we achieve superior perceptual quality. Experimental results demonstrate that STFTCodec outperforms both waveform-based and spectral-based approaches across multiple bitrates, while offering unique flexibility in compression ratio adjustment through STFT parameter modification without architectural changes.

音频压缩频域建模STFT神经编解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。