研究韩语语音问答中语音识别错误如何影响下游结果,发现微小错字可导致严重误解。
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

- 分析语音识别与大模型串联时错误传递机制,揭示信息损失主要来自识别阶段。
- 单字错误在韩语中尤为致命,极小误识即引发语义偏差导致回答失效。
- 直接输入音频的大型语音模型优于传统串行流程,适合噪声环境下的语音问答。
我们分析了自动语音识别(ASR)错误在韩语语音问答(SQA)中通过ASR--LLM级联传播的情况,重点关注传统ASR指标无法完全捕捉的下游语义失败。分析表明,不同性能的LLM在面对ASR错误时的相对下降程度一致,说明级联降级主要跟随ASR阶段的信息损失。进一步发现,韩语中单字错误是信息损失的关键来源,哪怕微小的转录差异也可能改变原意并降低下游问答性能。此外,辅助对比显示,在噪声韩语SQA任务中,具备近似语言主干的大规模音频语言模型表现优于传统ASR--LLM级联,表明直接音频输入有望缓解由转录引发的信息丢失。
原文摘要 · Abstract (English)
We analyze how automatic speech recognition (ASR) errors propagate through ASR--LLM cascades in Korean spoken question answering (SQA), focusing on downstream semantic failures that conventional ASR metrics cannot fully capture. Our analysis shows that the relative downstream degradation caused by ASR errors is consistent across LLMs with different absolute performance, suggesting that cascade degradation largely tracks ASR-stage information loss. We further identify single-character ASR errors as a particularly salient source of information loss in Korean, where even a minimal transcription difference can change the intended question and degrade downstream QA performance. Finally, an auxiliary comparison shows that a large audio language model outperforms an ASR--LLM cascade with an approximately matched language backbone in noisy Korean SQA, indicating the potential of direct audio input to mitigate transcript-induced information loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。