arXiv:2506.06732eess.AScs.AI2025-06中稿 · Interspeech 2025

用神经网络生成音频高频分量,提升编码效率与音质。

Neural Spectral Band Generation for Audio Coding

  • 用深度网络提取并重建高频分量的侧信息
  • 比HE-AAC-v1在更少侧信息下实现更好听感质量
  • 适合对音质和压缩率要求高的音频编码场景

频谱带复制(SBR)通过从低频带生成高频带实现高效编码,但仅依赖子带级的粗略频谱特征,难以适应多样化的声学信号。本文提出基于深度神经网络的生成方法——神经频谱带生成(n-SBG)。设计了一个端到端的编码器-解码器结构,用于提取并量化与高频成分相关的侧信息,并结合解码后的核心频带信号生成高频部分。整个编码流程采用生成对抗训练策略,以确保生成声音的感知真实性。在以AAC为核心编码器的实验中,所提方法在显著减少侧信息量的前提下,实现了优于HE-AAC-v1的主观音质。

原文摘要 · Abstract (English)

Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to diverse acoustic signals. In this paper, we explore the efficacy of a deep neural network (DNN)-based generative approach for coding the high-frequency bands, which we call neural spectral band generation (n-SBG). Specifically, we propose a DNN-based encoder-decoder structure to extract and quantize the side information related to the high-frequency components and generate the components given both the side information and the decoded core-band signals. The whole coding pipeline is optimized with generative adversarial criteria to enable the generation of perceptually plausible sound. From experiments using AAC as the core codec, we show that the proposed method achieves a better perceptual quality than HE-AAC-v1 with much less side information.

音频编码神经生成SBR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。