arXiv:2505.07235cs.SDeess.AS2025-05ICML被引 7

用心理声学指导频带重建,实现高保真音频压缩

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding

  • 按听觉敏感度分配码率,分频带量化提升压缩效率
  • 在多数据集上表现超越现有方法,高压缩比下仅损失12.5 Hz
  • 适合语音分离与生成任务,可对接语言模型使用

在神经音频编码中,如何在不同内容下保持高保真与感知质量仍是挑战。本文提出MUFFIN,一种全卷积神经心理声学编码框架,通过心理声学引导的多频带频谱重建实现高效压缩。核心是多频带谱残差向量量化(MBS-RVQ)模块,根据感知显著性动态分配码率,实现说话人身份与内容解耦。采用类Transformer卷积主干和改进蛇形激活函数,在细微频段增强分辨率。在多个基准测试中,MUFFIN持续优于现有方法;高压缩变体达12.5 Hz的最先进码率,感知损失极小。该模型在下游生成任务中也表现优异,证明其作为语言模型输入令牌表示的潜力。音频样本与代码已公开。

原文摘要 · Abstract (English)

Achieving high-fidelity audio compression while preserving perceptual quality across diverse content remains a key challenge in Neural Audio Coding (NAC). We introduce MUFFIN, a fully convolutional Neural Psychoacoustic Coding (NPC) framework that leverages psychoacoustically guided multi-band frequency reconstruction. At its core is a Multi-Band Spectral Residual Vector Quantization (MBS-RVQ) module that allocates bitrate across frequency bands based on perceptual salience. This design enables efficient compression while disentangling speaker identity from content using distinct codebooks. MUFFIN incorporates a transformer-inspired convolutional backbone and a modified snake activation to enhance resolution in fine-grained spectral regions. Experimental results on multiple benchmarks demonstrate that MUFFIN consistently outperforms existing approaches in reconstruction quality. A high-compression variant achieves a state-of-the-art 12.5 Hz rate with minimal loss. MUFFIN also proves effective in downstream generative tasks, highlighting its promise as a token representation for integration with language models. Audio samples and code are available.

音频编码心理声学向量量化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。