TQCodec是面向高保真音乐流媒体的神经音频编解码器,支持32-128kbps bitrate。
TQCodec: Towards neural audio codec for high-fidelity music streaming
- 基于SEANet架构,采用不平衡网络与SimVQ提升编码效率
- 在32-128kbps下实现优于现有方法的音质表现
- 适合高保真音乐流媒体、低延迟设备端部署
我们提出TQCodec,一种专为高比特率、高保真音乐流媒体设计的神经音频编解码器。与主要针对超低比特率(≤16kbps)的现有神经编解码器不同,TQCodec工作于44.1 kHz采样率,支持32 kbps至128 kbps的比特率,符合现代音乐流媒体平台的标准音质。模型采用基于SEANet的编码器-解码器架构,兼顾设备端高效计算,并引入多项改进:不平衡网络设计以低开销提升质量,SimVQ用于保留中频细节,以及相位感知波形损失。此外,我们提出一种感知驱动的带宽比特分配策略,优先保障感知关键的低频部分。在多种音乐数据集上的评估表明,TQCodec在目标比特率下实现了更优的音频质量,适用于高质量音频应用。
原文摘要 · Abstract (English)
We propose TQCodec, a neural audio codec designed for high-bitrate, high-fidelity music streaming. Unlike existing neural codecs that primarily target ultra-low bitrates (<= 16kbps), TQCodec operates at 44.1 kHz and supports bitrates from 32 kbps to 128 kbps, aligning with the standard quality of modern music streaming platforms. The model adopts an encoder-decoder architecture based on SEANet for efficient on-device computation and introduces several enhancements: an imbalanced network design for improved quality with low overhead, SimVQ for mid-frequency detail preservation, and a phase-aware waveform loss. Additionally, we introduce a perception-driven band-wise bit allocation strategy to prioritize perceptually critical lower frequencies. Evaluations on diverse music datasets demonstrate that TQCodec achieves superior audio quality at target bitrates, making it well-suited for high-quality audio applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。