arXiv:2411.19842eess.AScs.AI2024-11被引 90

用大规模Transformer和量化瓶颈,实现400-700比特/秒的高质量语音编码

Scaling Transformers for Low-Bitrate High-Quality Speech Coding

  • 采用大参数量Transformer配合可灵活调整的有限标量量化
  • 在400和700比特/秒下达到当前最优语音质量
  • 适合低带宽语音传输场景,如远程通信或嵌入式设备

神经音频编解码器模型对语音的分词是现代人工智能语音生成或理解流程中的关键环节,尤其在多模态背景下。传统方法倾向于使用参数量小、具有强先验偏置的组件。本文表明,通过将大规模参数量的Transformer架构应用于该任务,并结合灵活的有限标量量化(FSQ)瓶颈,可在极低比特率(400或700比特/秒)下实现最先进的语音质量。训练模型在客观与主观评测中均显著优于现有基线。

原文摘要 · Abstract (English)

The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with strong inductive biases. In this work we show that by scaling a transformer architecture with large parameter count to this problem, and applying a flexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible to reach state-of-the-art speech quality at extremely low bit-rates of $400$ or $700$ bits-per-second. The trained models strongly out-perform existing baselines in both objective and subjective tests.

语音编码Transformer量化低比特率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。