arXiv:2508.13376cs.CLcs.AI2025-08中稿 · IEEE ASRU 2025

用大模型知识蒸馏提升长语音转录的语法语义准确率

Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts

  • 用最优传输对齐音素与句向量,融合大模型语言上下文
  • 在Spoken Wikipedia上降低18.3%词错误率,实体识别准确率提升24%
  • 适合做长音频理解、信息提取的开发者与研究者参考

语音识别系统在长音频转录中常难以保持语法和语义准确性,影响命名实体识别(NER)、大小写和标点等任务。本文提出一种新方法,将LLaMA模型中的上下文知识蒸馏到Whisper中。采用两种策略:(1) 使用最优传输进行逐标记对齐,匹配维度与序列长度;(2) 最小化Whisper与LLaMA的句向量表示差异,融合语法与语义信息。在包含长音频与丰富实体的Spoken Wikipedia数据集上的评估显示,该方法显著提升词错误率(WER)、NER、大小写和标点成功率。通过引入新的NER评估指标并探索语义感知的语音识别,本工作强调了在转录中融入语言上下文的价值,为构建鲁棒的长时语音上下文感知识别系统奠定基础。

原文摘要 · Abstract (English)

ASR systems often struggle with maintaining syntactic and semantic accuracy in long audio transcripts, impacting tasks like Named Entity Recognition (NER), capitalization, and punctuation. We propose a novel approach that enhances ASR by distilling contextual knowledge from LLaMA models into Whisper. Our method uses two strategies: (1) token level distillation with optimal transport to align dimensions and sequence lengths, and (2) representation loss minimization between sentence embeddings of Whisper and LLaMA, blending syntax and semantics. Evaluations on the Spoken Wikipedia dataset, a benchmark with long audios and rich entities demonstrate significant improvements in Word Error Rate (WER), NER, capitalization, and punctuation success. By introducing novel NER metrics and exploring semantics aware ASR, our work highlights the value of integrating linguistic context into transcription, setting a foundation for robust, context-aware ASR in longform speech.

语音识别知识蒸馏长语音语义感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。