用焦点调制网络实现0.16-0.65kbps低比特语音编码
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
- 基于焦点调制机制,单二值码本实现高效压缩
- 在0.16~0.65kbps下优于现有最先进方法
- 支持多语言、噪声环境,适合生成建模
大型语言模型通过大规模数据上的自监督预训练彻底改变了自然语言处理。受此启发,研究者尝试将类似方法应用于语音,通过神经音频编码器将连续音频离散化为词元。然而,现有方法存在比特率高、语义或声学信息丢失、需多码本设计以同时保留两者等问题,增加了下游任务的架构复杂性。为此,我们提出FocalCodec,一种基于焦点调制的高效低比特率编码器,采用单个二值码本,在0.16至0.65 kbps范围内压缩语音。FocalCodec在语音重合成与语音转换任务中表现优异,比特率低于当前最优水平,同时能有效处理多语言语音与噪声环境。下游任务评估表明,FocalCodec成功保留了充分的语义与声学信息,且适用于生成建模。演示样本与代码已公开于 https://lucadellalib.github.io/focalcodec-web/。
原文摘要 · Abstract (English)
Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discretizing continuous audio into tokens using neural audio codecs. However, existing approaches face limitations, including high bitrates, the loss of either semantic or acoustic information, and the reliance on multi-codebook designs when trying to capture both, which increases architectural complexity for downstream tasks. To address these challenges, we introduce FocalCodec, an efficient low-bitrate codec based on focal modulation that utilizes a single binary codebook to compress speech between 0.16 and 0.65 kbps. FocalCodec delivers competitive performance in speech resynthesis and voice conversion at lower bitrates than the current state-of-the-art, while effectively handling multilingual speech and noisy environments. Evaluation on downstream tasks shows that FocalCodec successfully preserves sufficient semantic and acoustic information, while also being well-suited for generative modeling. Demo samples and code are available at https://lucadellalib.github.io/focalcodec-web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。