用大模型生成可控语音错误,提升对话系统在听错情况下的准确性
Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking
- 用关键词提示控制错误位置,生成语音相似的错误
- 在低准确率语音识别环境下,对话状态跟踪准确率显著提升
- 适合研究语音交互鲁棒性或数据增强的工程师
对话状态跟踪(DST)是任务导向对话系统的关键组件,用于识别对话中的关键信息。然而,在语音对话环境中,由于自动语音识别(ASR)系统的命名实体识别错误,DST性能显著下降。本文提出一种简单但有效的新数据增强方法,针对这些实体进行优化,以提升DST模型的鲁棒性。该方法通过关键词标注提示,精确控制错误插入位置,并引入语音上相近的错误模式。实验表明,该方法在关键词上生成了充分的错误模式,在噪声环境和低准确率ASR条件下均提升了模型表现。
原文摘要 · Abstract (English)
Dialogue State Tracking (DST) is a key part of task-oriented dialogue systems, identifying important information in conversations. However, its accuracy drops significantly in spoken dialogue environments due to named entity errors from Automatic Speech Recognition (ASR) systems. We introduce a simple yet effective data augmentation method that targets those entities to improve the robustness of DST model. Our novel method can control the placement of errors using keyword-highlighted prompts while introducing phonetically similar errors. As a result, our method generated sufficient error patterns on keywords, leading to improved accuracy in noised and low-accuracy ASR environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。