arXiv:2606.11279eess.AScs.CL2026-06中稿 · Interspeech 2026

用极小内存实现海量关键词的快速语音检索

Massive Open-Vocabulary Keyword Spotting

论文配图:Massive Open-Vocabulary Keyword Spotting
图 1 · 摘自论文原文
  • 采用压缩特征存储技术,内存占用仅为基线1/128
  • 支持百万级关键词库,识别准确率接近无损方案
  • 无需微调模型,跨语言关键词检索仍有效

自动语音识别系统在处理训练数据中罕见的专业术语时表现不佳。开放词汇关键词检测结合上下文偏置可缓解此问题,但现有系统仅能处理数百个关键词,否则成为性能瓶颈。我们提出一种新系统,将特征存储内存降低至可比基线的1/128,可在不微调语音识别模型的前提下,支持大规模关键词库的开放词汇检索,即使在训练中未见的语言上也能达到与无损方案相当的实体召回率。

原文摘要 · Abstract (English)

Automatic speech recognition systems have been shown to under-perform when it comes to transcribing words rarely seen in the training data, namely specialized terminology. Open-vocabulary keyword spotting, combined with contextual biasing, has been shown to mitigate this issue. However, existing systems can only handle glossaries of a few hundred terms without becoming an infeasible bottleneck. We propose a system that stores features with a memory footprint up to 128 times smaller than a comparable baseline and allows users to process massive databases while remaining open-vocabulary. Without fine-tuning the speech recognition model, our system achieves a comparable entity recall as uncompressed solutions, even in languages not seen during training.

语音识别关键词检测内存优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。