解决语音识别中大量关键词识别率下降的问题
H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
- 通过声学相似度匹配提前筛选关键词候选
- 在传统模型上提升关键词召回率,效果显著
- 适配主流语音识别系统,易集成且通用性强
关键词定制对提升特定领域术语识别准确率至关重要。尽管传统模型和音频大语言模型(Audio LLMs)推动了该方向发展,但现有方法在处理大规模关键词时性能急剧下降。本文提出一种可插拔的关键词预检索模块(H-PRM),通过计算关键词与语音片段之间的声学相似度,高效筛选最相关的关键词候选。该方案可无缝集成至SeACo-Paraformer等传统模型,显著提升关键词后召回率(PRR)。同时,采用基于提示的方法将H-PRM引入Audio LLMs,实现无损关键词定制。大量实验验证其优于现有方法,为语音识别中的关键词定制提供了新路径。
原文摘要 · Abstract (English)
Hotword customization is crucial in ASR to enhance the accuracy of domain-specific terms. It has been primarily driven by the advancements in traditional models and Audio large language models (LLMs). However, existing models often struggle with large-scale hotwords, as the recognition rate drops dramatically with the number of hotwords increasing. In this paper, we introduce a novel hotword customization system that utilizes a hotword pre-retrieval module (H-PRM) to identify the most relevant hotword candidate by measuring the acoustic similarity between the hotwords and the speech segment. This plug-and-play solution can be easily integrated into traditional models such as SeACo-Paraformer, significantly enhancing hotwords post-recall rate (PRR). Additionally, we incorporate H-PRM into Audio LLMs through a prompt-based approach, enabling seamless customization of hotwords. Extensive testing validates that H-PRM can outperform existing methods, showing a new direction for hotword customization in ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。