arXiv:2410.15609cs.CLcs.SD2024-10EMNLP被引 1

让语音理解模型更抗错,用通用噪声训练提升跨系统泛化能力

Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding

  • 通过切断噪声的非因果影响,生成适用于任何语音识别系统的普适噪声
  • 在多个未见语音识别系统上测试,模型鲁棒性显著提升
  • 适合需要跨平台部署的语音理解系统开发者

近年来,预训练语言模型(PLMs)被广泛用于语音理解(SLU)。然而,自动语音识别(ASR)系统常产生错误转录,导致语音理解模型输入存在噪声,严重影响其性能。为解决此问题,我们旨在通过引入常见于ASR系统的合理噪声,使SLU模型具备抵抗ASR错误的能力。现有语音噪声注入(SNI)方法虽能实现此目标,但存在对特定ASR系统偏倚的问题。本文提出一种新方法,通过消除噪声的非因果效应,生成对任意ASR系统都合理的噪声。实验与分析表明,该方法能在提前引入更丰富且合理的ASR噪声的基础上,有效提升SLU模型对未见ASR系统的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Recently, pre-trained language models (PLMs) have been increasingly adopted in spoken language understanding (SLU). However, automatic speech recognition (ASR) systems frequently produce inaccurate transcriptions, leading to noisy inputs for SLU models, which can significantly degrade their performance. To address this, our objective is to train SLU models to withstand ASR errors by exposing them to noises commonly observed in ASR systems, referred to as ASR-plausible noises. Speech noise injection (SNI) methods have pursued this objective by introducing ASR-plausible noises, but we argue that these methods are inherently biased towards specific ASR systems, or ASR-specific noises. In this work, we propose a novel and less biased augmentation method of introducing the noises that are plausible to any ASR system, by cutting off the non-causal effect of noises. Experimental results and analyses demonstrate the effectiveness of our proposed methods in enhancing the robustness and generalizability of SLU models against unseen ASR systems by introducing more diverse and plausible ASR noises in advance.

语音理解噪声注入鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。