arXiv:2505.14814cs.SDcs.CL2025-05中稿 · Interspeech 2025被引 2

通过修改关键词字符生成难例,显著提升语音关键词检测模型性能。

GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples

  • 对关键词字符进行增删替换,系统生成靠近分类边界的难例。
  • 在合成难例数据集上AUC提升61%,且不影响正例和普通负例表现。
  • 适合提升语音识别中关键词检测的鲁棒性,尤其适用于边界案例不足场景。

语音关键词检测(KWS)旨在判断音频中是否存在特定关键词。模型性能取决于其对关键词与非关键词边界附近样本的判别能力。然而,这类边界样本在训练数据中通常稀缺,制约了模型表现。本文提出一种系统化方法,通过对关键词的字符(graphemes)进行插入、删除或替换操作,生成接近决策边界的对抗性样本。在某主流关键词的保留测试数据上评估该方法,结果表明,在合成难例数据集上,模型AUC提升了61%,同时保持了对正例及环境负例音频的质量表现。

原文摘要 · Abstract (English)

Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data.

关键词检测语音识别对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。