arXiv:2512.00511eess.AS2025-12中稿 · 2026 Data Compress…

用参数化加噪提升语音压缩后识别效果,1比特下错误率降25%。

A Low-Complexity Speech Codec Using Parametric Dithering for ASR

  • 用可调参数的加噪技术优化语音压缩,让识别更鲁棒。
  • 1比特时相对识别错误率降低25%,2/3比特分别降32.4%和33.5%。
  • 适合低码率语音传输场景,兼顾性能与压缩效率。

加噪是一种常用于提升有损数据压缩感知质量的技术。本文从理论和实验两方面证明了在语音识别(ASR)输入压缩中使用加噪的合理性。我们建立了在有损输入压缩下最优ASR性能的分析框架,并据此提出一种适用于低复杂度语音压缩流水线的参数化加噪方法。该方法在1比特分辨率下实现25%的相对词错误率(CER)改进,在2比特和3比特分辨率下分别取得32.4%和33.5%的改进;第二种加噪策略进一步降低了数据率。所提编码器可灵活适配性能目标或熵约束。

原文摘要 · Abstract (English)

Dithering is a technique commonly used to improve the perceptual quality of lossy data compression. In this work, we analytically and experimentally justify the use of dithering for ASR input compression. We formalize an understanding of optimal ASR performance under lossy input compression and leverage this to propose a parametric dithering technique for a low-complexity speech compression pipeline. The method performs well at 1-bit resolution, showing a 25\% relative CER improvement, while also demonstrating improvements of 32.4\% and 33.5\% at 2- and 3-bit resolution, respectively, with our second dither choice yielding a reduced data rate. The proposed codec is adaptable to meet performance targets or stay within entropy constraints.

语音压缩语音识别加噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。