评估大模型在语音问答中识别问题无法回答并主动修复对话的能力。
Pardon? Evaluating Conversational Repair in Large Audio-Language Models
- 设计语义-声学掩码协议,区分可回答与不可回答的语音输入。
- 提出EAR分数,同时衡量回答准确性和未答时的修复行为。
- 发现多数模型能答对问题却不会识别无法回答并请求澄清。
大型音频-语言模型(LALMs)在语音问答任务中表现优异,现有评估多关注答案准确性和对声学干扰的鲁棒性。但这些评估隐含假设:语音输入始终具备可回答性,这在真实交互中常不成立,因关键信息缺失导致问题无法回答。本文提出一种修复感知的评估设置,明确区分可回答与不可回答的音频输入。将可回答性定义为输入本身的属性,并采用语义-声学掩码协议构建配对评估条件。基于此,提出评价能力意识与修复(EAR)分数,一种非补偿性指标,联合评估在可回答条件下的任务能力与在不可回答条件下的修复行为。在两个语音问答基准上对多种LALMs的实验表明,答案准确率与对话可靠性之间存在持续差距:尽管许多模型在输入可回答时表现良好,但多数无法识别语义不可回答性,也未能启动适当的对话修复。研究揭示了当前以准确率为中心的评估范式的局限性,呼吁将不可回答输入视为修复提示,推动更可靠的交互评估。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs) have demonstrated strong performance in spoken question answering (QA), with existing evaluations primarily focusing on answer accuracy and robustness to acoustic perturbations. However, such evaluations implicitly assume that spoken inputs remain semantically answerable, an assumption that often fails in real-world interaction when essential information is missing. In this work, we introduce a repair-aware evaluation setting that explicitly distinguishes between answerable and unanswerable audio inputs. We define answerability as a property of the input itself and construct paired evaluation conditions using a semantic-acoustic masking protocol. Based on this setting, we propose the Evaluability Awareness and Repair (EAR) score, a non-compensatory metric that jointly evaluates task competence under answerable conditions and repair behavior under unanswerable conditions. Experiments on two spoken QA benchmarks across diverse LALMs reveal a consistent gap between answer accuracy and conversational reliability: while many models perform well when inputs are answerable, most fail to recognize semantic unanswerability and initiate appropriate conversational repair. These findings expose a limitation of prevailing accuracy-centric evaluation practices and motivate reliability assessments that treat unanswerable inputs as cues for repair and continued interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。