arXiv:2604.14654cs.SDeess.AS2026-04

用强化学习优化200bps语音编码的可懂度,提升通信清晰度。

ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning

  • 将量化建模为随机策略,用强化学习直接优化可懂度
  • 200bps下测试集词错误率降至3.20%,相对降低13%
  • 适合卫星、水下等超低带宽通信场景使用

在卫星和水下等带宽受限的通信场景中,语音常需以极低比特率传输,此时可懂度是首要目标。传统基于声学重建损失训练的编解码器会将比特分配给感知细节,导致词错误率(WER)显著上升。本文提出ClariCodec,一种运行在200 bit per second(bps)的神经语音编解码器,将量化过程重构为随机策略,实现基于强化学习(RL)的可懂度优化。具体而言,编码器在保持声学重建路径冻结的前提下,使用基于WER的奖励进行微调。即使不启用强化学习,ClariCodec在LibriSpeech test-clean集上已达到3.68%的WER,表现媲美更高比特率的编解码器。进一步通过强化学习微调后,test-clean集上WER降至3.20%,test-other集为8.93%,相对减少13%,同时保持良好听觉质量。

原文摘要 · Abstract (English)

In bandwidth-constrained communication such as satellite and underwater channels, speech must often be transmitted at ultra-low bitrates where intelligibility is the primary objective. At such extreme compression levels, codecs trained with acoustic reconstruction losses tend to allocate bits to perceptual detail, leading to substantial degradation in word error rate (WER). This paper proposes ClariCodec, a neural speech codec operating at 200 bit per second (bps) that reformulates quantisation as a stochastic policy, enabling reinforcement learning (RL)-based optimisation of intelligibility. Specifically, the encoder is fine-tuned using WER-driven rewards while the acoustic reconstruction pipeline remains frozen. Even without RL, ClariCodec achieves 3.68% WER on the LibriSpeech test-clean set at 200 bps, already competitive with codecs operating at higher bitrates. Further RL fine-tuning reduces WER to 3.20% on test-clean and 8.93% on test-other, corresponding to a 13% relative reduction while preserving perceptual quality.

语音编码强化学习低比特率可懂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。