用少量离散符号实现接近原始音质的语音编码
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
- 基于RVQGAN模型,用多源语音数据微调出高效语音分词器
- 每秒仅需150-300个令牌(1500-3000 bps)即可实现近乎无损重建
- 适合对低码率语音压缩有需求的研究与工业应用
离散音频编解码器因大语言模型能学习其压缩声学表征而重新受到关注。尽管已有可训练的离散分词器取得显著成果,但多数需高令牌率才能保证高质量重建。本研究利用多样化的开源语音数据,微调了一个开源通用音频RVQGAN模型,涵盖不同录音条件与质量水平。所得到的宽带(24kHz)专用语音模型,在每秒150-300个令牌(1500-3000 bps)的速率下,实现几乎与脉冲编码调制(PCM)无法区分的语音重建。评估使用了涵盖多种录音环境的英文语音数据集,包括录音室场景。语音样本已公开于 http://ibm.biz/IS24SpeechRVQ,模型正式发布于 https://huggingface.co/ibm/DAC.speech.v1.0。
原文摘要 · Abstract (English)
Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. Various publicly available trainable discrete tokenizers recently demonstrated impressive results for audio tokenization, yet they mostly require high token rates to gain high-quality reconstruction. In this study, we fine-tuned an open-source general audio RVQGAN model using diverse open-source speech data, considering various recording conditions and quality levels. The resulting wideband (24kHz) speech-only model achieves speech reconstruction, which is nearly indistinguishable from PCM (pulse-code modulation) with a rate of 150-300 tokens per second (1500-3000 bps). The evaluation used comprehensive English speech data encompassing different recording conditions, including studio settings. Speech samples are made publicly available in http://ibm.biz/IS24SpeechRVQ . The model is officially released in https://huggingface.co/ibm/DAC.speech.v1.0
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。