用0.275kbps实现高保真音频压缩,突破传统编码极限。
High-Fidelity Generative Audio Compression at 0.275kbps
- 以任务有效性替代波形重建,利用生成模型分担传输负担。
- 在0.275kbps下实现32kHz音频高保真还原,压缩比达3000倍。
- 适合低带宽通信与语音-语言生成等需要语义保真的场景。
在超低比特率下实现高质量通用音频压缩对低带宽通信和生成式音频-语言建模至关重要。传统音频压缩方法与现代神经编解码器本质上均以波形重建为目标,因此在极低比特率下性能急剧下降,常导致严重声学伪影和显著语义失真。为克服此局限,本文提出生成式音频压缩(GAC),从信号保真度转向任务导向的有效性,基于信息容量定律,在接收端利用强大计算能力补偿通信瓶颈,体现‘多计算,少带宽’理念。通过在发送端融合语义理解、接收端采用可扩展生成合成,将信息负担交由强大的模型先验处理。所提1.8B参数模型在0.275kbps下实现了32kHz通用音频的高保真重建,即使在0.175kbps下仍具备良好可懂性,压缩比约达3000倍,显著优于现有最先进神经编解码器,在感知质量和语义一致性上表现更优。
原文摘要 · Abstract (English)
High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditional audio compression methods and contemporary neural codecs are fundamentally designed for waveform reconstruction. As a result, when operating at ultra-low bitrates, these methods degrade rapidly and often fail to preserve essential information, leading to severe acoustic artifacts and pronounced semantic distortion. To overcome these limitations, we introduce Generative Audio Compression (GAC), a novel paradigm shift from signal fidelity to task-oriented effectiveness. Implemented within the AI Flow framework, GAC is theoretically grounded in the Law of Information Capacity. These foundations posit that abundant computational power can be leveraged at the receiver to offset extreme communication bottlenecks--exemplifying the More Computation, Less Bandwidth philosophy. By integrating semantic understanding at the transmitter with scalable generative synthesis at the receiver, GAC offloads the information burden to powerful model priors. Our 1.8B-parameter model achieves high-fidelity reconstruction of 32kHz general audio at an unprecedented bitrate of 0.275kbps. Even at 0.175kbps, it still preserves a strong intelligible audio transmission capability, which represents an about 3000x compression ratio, significantly outperforming current state-of-the-art neural codecs in maintaining both perceptual quality and semantic consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。