构建音频问答不可答性评估基准,提升模型可信度
AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
- 设计三类不可答场景:答案缺失、选项不匹配、问题与音频无关
- 实测表明模型在可答任务表现好,但在不可答场景严重失准
- 适合研究音频-语言模型鲁棒性与可信性的学者使用
近期音频感知大模型在音频问答任务中表现出色。然而,现有评测基准主要集中于可答问题,忽略了现实中常见但未被充分关注的不可答情况——即从音频中无法可靠推断出答案的情形。此类问题常因误导性、表述不清或信息不匹配而出现。为此,我们提出AQUA-Bench,一个面向音频问答不可答性评估的基准。该基准系统评估三种典型场景:答案缺失检测(正确选项不存在)、答案集不兼容检测(选项与问题类别不符)、音频-问题不兼容检测(问题与音频无关或缺乏充分依据)。通过量化这些场景的表现,AQUA-Bench为模型可靠性提供了严格衡量标准,推动音频-语言系统向更稳健、更可信方向发展。实验显示,尽管模型在标准可答任务上表现优异,但在不可答情形下仍面临显著挑战,暴露出当前音频-语言理解中的盲区。
原文摘要 · Abstract (English)
Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks mainly cover answerable questions and overlook the challenge of unanswerable ones, where no reliable answer can be inferred from the audio. Such cases are common in real-world settings, where questions may be misleading, ill-posed, or incompatible with the information. To address this gap, we present AQUA-Bench, a benchmark for Audio Question Unanswerability Assessment. It systematically evaluates three scenarios: Absent Answer Detection (the correct option is missing), Incompatible Answer Set Detection (choices are categorically mismatched with the question), and Incompatible Audio Question Detection (the question is irrelevant or lacks sufficient grounding in the audio). By assessing these cases, AQUA-Bench offers a rigorous measure of model reliability and promotes the development of audio-language systems that are more robust and trustworthy. Our experiments suggest that while models excel on standard answerable tasks, they often face notable challenges with unanswerable ones, pointing to a blind spot in current audio-language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。