arXiv:2506.01263cs.CLcs.SD2025-06中稿 · Interspeech 2025被引 1

不重训练即可提升语音识别对生僻词的准确率。

WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing

  • 用中间层声学特征做关键词检测,动态加偏置修正输出
  • 日本语数据集上未知词F1提升29%
  • 无需重训练,适合已部署的大模型快速优化

尽管端到端语音识别技术取得进展,其输出仍倾向于训练数据中的词汇,导致专有名词等生僻词识别不准。为此,我们提出一种无需重训练的方法,通过在推理阶段利用中间层声学特征进行关键词检测,并对后续层施加偏置,提升罕见词识别精度。关键词检测采用快速且容忍模糊匹配的通配符CTC,可灵活处理难以严格匹配的词汇。该方法无需重新训练现有模型,适用于大规模模型。在日本语语音识别实验中,未知词的F1分数提升了29%。

原文摘要 · Abstract (English)

Despite recent advances in end-to-end speech recognition methods, the output tends to be biased to the training data's vocabulary, resulting in inaccurate recognition of proper nouns and other unknown terms. To address this issue, we propose a method to improve recognition accuracy of such rare words in CTC-based models without additional training or text-to-speech systems. Specifically, keyword spotting is performed using acoustic features of intermediate layers during inference, and a bias is applied to the subsequent layers of the acoustic model for detected keywords. For keyword detection, we adopt a wildcard CTC that is both fast and tolerant of ambiguous matches, allowing flexible handling of words that are difficult to match strictly. Since this method does not require retraining of existing models, it can be easily applied to even large-scale models. In experiments on Japanese speech recognition, the proposed method achieved a 29% improvement in the F1 score for unknown words.

语音识别关键词检测无重训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。