用可微分听觉感知损失提升音频增强效率与质量
Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss

- 用Mamba替代注意力和LSTM,降低计算开销
- 推理速度提升14倍,内存仅用1/5,主观评分高15%
- 针对压缩音频优化,修复MP3后质量提升52%
音频增强旨在提升音频的感知质量。本文提出AEROMamba_{P},作为AERO超分辨率架构的高效变体,将注意力和LSTM层替换为Mamba状态空间模型,并引入基于感知音频质量度量(PAQM)的新可微分感知损失。训练时所需GPU内存仅为基线的2-4倍;推理时速度提升14倍,仅使用1/5的GPU内存。在将钢琴数据集和MUSDB18从11.025 kHz上采样至44.1 kHz的主观听感测试中,其感知质量得分比AERO高出15%。进一步提出AEROMamba_{PS},采用相同框架但以PAQM损失替代STFT重建损失,专门用于修复32 kbps MP3编码音频。听感评估显示,该方法在恢复压缩音频时的质量评分比AEROMamba_{P}高52%。结果表明,结合PAQM驱动训练与轻量级状态空间建模,可在带宽受限和压缩音频场景下同时实现高感知质量与计算效率。
原文摘要 · Abstract (English)
Audio enhancement consists of improving the perceived quality of audio signals. Initially, with the aim of addressing bandwidth extension, this work proposes \(AEROMamba_{P}\), an efficient variant of the AERO super-resolution architecture where attention and LSTM layers are replaced by the Mamba state-space model, and which incorporates a newly developed differentiable perceptual loss derived from the Perceptual Audio Quality Measure (PAQM). During training, the architecture requires approximately 2-4x less GPU memory than the baseline; during inference, it achieves a 14x speedup while using only one-fifth of the GPU memory. When upsampling both a piano dataset and MUSDB18 from 11.025 kHz to 44.1 kHz, subjective listening tests show that \(AEROMamba_{P}\) outperforms AERO by 15% in perceived quality scores. Next, to handle the enhancement of audio signals that have been highly compressed by lossy coding, it is further proposed \(AEROMamba_{PS}\), which applies the same framework but replaces STFT reconstruction losses with the PAQM loss, specifically to enhance MP3 encoded audio at 32 kbps. In listening evaluations, \(AEROMamba_{PS}\) achieves 52% higher quality rating than \(AEROMamba_{P}\) when restoring compressed audio. These results demonstrate that PAQM-driven training coupled with lightweight state-space modeling yields high perceptual quality and computational efficiency in both band-limited and compressed audio scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。