arXiv:2410.18363cs.AIcs.SD2024-10ICML被引 7

不微调模型,用上下文引导提升专业词汇识别准确率

Contextual Biasing to Improve Domain-specific Custom Vocabulary Audio Transcription without Explicit Fine-Tuning of Whisper Model

  • 用神经符号前缀树引导模型输出特定词汇
  • 在航海数据上词错误率显著降低
  • 适合缺乏标注数据的专业领域应用

OpenAI的Whisper语音识别模型在跨数据集和领域上具有出色的泛化能力,但其广泛适应性可能导致对特定词汇识别性能下降。传统解决方案需微调模型,依赖大量标注音频数据,而这些数据在特定领域往往难以获取。本文提出一种无需显式微调或修改模型参数的方法,仅使用少量训练数据,通过上下文引导(contextual biasing)机制,结合神经符号前缀树结构,引导Whisper模型输出偏向特定词汇。在模拟训练环境中收集的航海数据验证集上,对比不同参数规模的原始Whisper模型与我们的偏置模型,结果显示词错误率明显下降,下游应用性能显著提升。研究表明,该方法在词汇量有限的领域中具有改善语音转写性能的潜力。

原文摘要 · Abstract (English)

OpenAI's Whisper Automated Speech Recognition model excels in generalizing across diverse datasets and domains. However, this broad adaptability can lead to diminished performance in tasks requiring recognition of specific vocabularies. Addressing this challenge typically involves fine-tuning the model, which demands extensive labeled audio data that is often difficult to acquire and unavailable for specific domains. In this study, we propose a method to enhance transcription accuracy without explicit fine-tuning or altering model parameters, using a relatively small training dataset. Our method leverages contextual biasing, to direct Whisper model's output towards a specific vocabulary by integrating a neural-symbolic prefix tree structure to guide the model's transcription output. To validate our approach, we conducted experiments using a validation dataset comprising maritime data collected within a simulated training environment. A comparison between the original Whisper models of varying parameter sizes and our biased model revealed a notable reduction in transcription word error rate and enhanced performance of downstream applications. Our findings suggest that this methodology holds promise for improving speech-to-text translation performance in domains characterized by limited vocabularies.

语音识别上下文引导Whisper小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。