arXiv:2409.06263cs.CLcs.AI2024-09中稿 · AACL-IJCNLP 2025被引 1

用大模型生成可控语音错误,提升对话系统在听错情况下的准确性

Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking

  • 用关键词提示控制错误位置,生成语音相似的错误
  • 在低准确率语音识别环境下,对话状态跟踪准确率显著提升
  • 适合研究语音交互鲁棒性或数据增强的工程师

对话状态跟踪(DST)是任务导向对话系统的关键组件,用于识别对话中的关键信息。然而,在语音对话环境中,由于自动语音识别(ASR)系统的命名实体识别错误,DST性能显著下降。本文提出一种简单但有效的新数据增强方法,针对这些实体进行优化,以提升DST模型的鲁棒性。该方法通过关键词标注提示,精确控制错误插入位置,并引入语音上相近的错误模式。实验表明,该方法在关键词上生成了充分的错误模式,在噪声环境和低准确率ASR条件下均提升了模型表现。

原文摘要 · Abstract (English)

Dialogue State Tracking (DST) is a key part of task-oriented dialogue systems, identifying important information in conversations. However, its accuracy drops significantly in spoken dialogue environments due to named entity errors from Automatic Speech Recognition (ASR) systems. We introduce a simple yet effective data augmentation method that targets those entities to improve the robustness of DST model. Our novel method can control the placement of errors using keyword-highlighted prompts while introducing phonetically similar errors. As a result, our method generated sufficient error patterns on keywords, leading to improved accuracy in noised and low-accuracy ASR environments.

对话系统语音识别数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。