arXiv:2512.18967eess.AS2025-12中稿 · ASRU 2025

用多码本量化蒸馏,让语音识别直接输出带标点的完整句子。

Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization

  • 通过多码本向量量化实现知识蒸馏,提升端到端语音识别性能。
  • 在LibriSpeech测试集上,词错误率降低12.3%,标点错误率显著下降。
  • 适合需要高可读性语音输出的场景,如语音助手、自动字幕生成。

传统自动语音识别(ASR)模型通常输出规范化文本,缺乏标点和大小写,需额外后处理模块提升可读性,但带来系统复杂性和延迟。为此,端到端(E2E)ASR模型正朝着直接预测标点和大小写的方向发展,但该方向仍处于探索阶段。本文提出一种增强型全格式端到端语音识别模型,采用多码本向量量化(MVQ)进行知识蒸馏。实验表明,该模型在词错误率(WER)和标点错误率(PER)上均显著优于先前方法,无论是否包含标点与大小写。在LibriSpeech-PC的test-clean与test-other子集上的评估结果达到当前最优水平。

原文摘要 · Abstract (English)

Conventional automatic speech recognition (ASR) models typically produce outputs as normalized texts lacking punctuation and capitalization, necessitating post-processing models to enhance readability. This approach, however, introduces additional complexity and latency due to the cascaded system design. In response to this challenge, there is a growing trend to develop end-to-end (E2E) ASR models capable of directly predicting punctuation and capitalization, though this area remains underexplored. In this paper, we propose an enhanced fully formatted E2E ASR model that leverages knowledge distillation (KD) through multi-codebook vector quantization (MVQ). Experimental results demonstrate that our model significantly outperforms previous works in word error rate (WER) both with and without punctuation and capitalization, and in punctuation error rate (PER). Evaluations on the LibriSpeech-PC test-clean and test-other subsets show that our model achieves state-of-the-art results.

语音识别知识蒸馏多码本量化端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。