分频段压缩音频,让语音音乐都保真。
BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
- 将音频频谱拆成独立频段分别压缩,适配不同内容特性。
- 在语音、音乐和音效上均优于基线模型,全场景表现均衡。
- 适合需要通用高质量音频重建的场景,如多类型音频处理。
神经音频编解码器近年来在高压缩率下实现了高保真语音重建。但语音与非语音音频具有根本不同的频谱特征:语音能量集中在基频谐波附近(80-400 Hz)的窄带区域,而非语音音频需在整个频谱范围内忠实还原,尤其要保留影响音色和质感的高频成分。这导致专为语音优化的编解码器在音乐或声音上表现退化。整体处理全频带的方式效率低下:不同频段的信息密度和听觉重要性随内容类型差异显著,而全带方法对所有频率分配相同容量,未考虑声学结构。为此,我们提出BSCodec(Band-Split Codec),一种新型神经音频编解码架构,将频谱维度拆分为多个独立频段,分别进行压缩。实验表明,当在语音、音乐和声音的混合数据集上训练时,BSCodec在音效和音乐上的重建质量显著优于基线模型,同时在语音领域仍保持竞争力。下游任务评估进一步验证其在实际应用中的潜力。
原文摘要 · Abstract (English)
Neural audio codecs have recently enabled high-fidelity reconstruction at high compression rates, especially for speech. However, speech and non-speech audio exhibit fundamentally different spectral characteristics: speech energy concentrates in narrow bands around pitch harmonics (80-400 Hz), while non-speech audio requires faithful reproduction across the full spectrum, particularly preserving higher frequencies that define timbre and texture. This poses a challenge: speech-optimized neural codecs suffer degradation on music or sound. Treating the full spectrum holistically is suboptimal: frequency bands have vastly different information density and perceptual importance by content type, yet full-band approaches apply uniform capacity across frequencies without accounting for these acoustic structures. To address this gap, we propose BSCodec (Band-Split Codec), a novel neural audio codec architecture that splits the spectral dimension into separate bands and compresses each band independently. Experimental results demonstrate that BSCodec achieves superior reconstruction over baselines across sound and music, while maintaining competitive quality in the speech domain, when trained on the same combined dataset of speech, music and sound. Downstream benchmark tasks further confirm that BSCodec shows strong potential for use in downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。