arXiv:2605.25404cs.CLeess.AS2026-05

通过细粒度错误诊断,提升语音对话系统在噪声和口音下的鲁棒性。

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

论文配图:Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
图 1 · 摘自论文原文
  • 基于ASR隐层特征,区分感知、理解与删除三类错误
  • 域迁移场景下错误召回率提升至57.96%(基线23.66%)
  • 减少30%字错误率,下游任务提升17%,适合高可靠性场景

级联式自动语音识别-大语言模型(ASR-LLM)管道在工业语音对话系统中仍占主导地位,因其解耦设计可保证感知可验证性。然而,级联系统存在误差传播问题,转录失败会传递至后续模块,降低交互质量。尽管ASR置信度分数可简单过滤不可靠输入,但该方法根本局限在于无法检测删除错误,也无法区分听不清(感知)与听不懂(理解)的差异,而这两类错误需不同恢复策略。本文提出一种原因感知的错误恢复范式,重新思考语音对话系统的鲁棒性。不同于传统置信度过滤,我们引入一系列小型高精度探测器,利用深度ASR隐层表示,将词级别错误分解为感知、理解与删除三类失败。这种细粒度诊断能力使大语言模型能协调针对性的多轮澄清策略,有效将模糊信号转化为流畅交互。实验验证了该方法的精度:在域迁移错误上的召回率超过基线两倍(57.96% vs. 23.66%)。关键的是,该诊断精度带来最高30%的字错误率(WER)下降,以及在多种口音、失真和领域下下游任务17%的性能提升。

原文摘要 · Abstract (English)

Cascaded Automatic Speech Recognition -- Large Language Model (ASR-LLM) pipelines remain popular for industrial Spoken Dialogue Systems (SDS), primarily because their decoupled design ensures perceptual verifiability. However, cascaded systems suffer from error propagation, as transcription failures inevitably cascade to subsequent components, thereby degrading the final interaction quality. Although ASR confidence scores offer a simple filter for unreliable inputs, this approach is fundamentally limited because it typically fails to detect deletion errors or to distinguish between acoustic (inability to hear clearly) and linguistic (inability to understand) mismatches, both of which require targeted recovery strategies. In this paper, we propose a cause-aware error recovery paradigm that fundamentally rethinks robustness in SDS. Unlike traditional confidence filtering, we introduce a suite of small precision-focused detectors that exploit deep ASR latent representations to disentangle token-level errors into perception, comprehension, and deletion failures. This fine-grained diagnostic intelligence empowers the LLM to orchestrate targeted, multi-turn clarification strategies, effectively transforming ambiguous signals into seamless user interactions. Experimental results validate the precision of our approach, which more than doubles the recall on domain-shift errors (57.96% vs. 23.66%) compared to baselines. Crucially, this diagnostic precision yields up to a 30% reduction in WER and a 17% improvement on the downstream task across diverse accents, distortions, and domains.

语音对话错误诊断鲁棒性ASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。