arXiv:2510.04934eess.AScs.AI2025-10

提出AURA评分,更准确评估语音问答模型的开放回答质量。

AURA Score: A Metric For Holistic Audio Question Answering Evaluation

  • 构建首个1万条人工标注的语音问答基准集AQEval,支持系统性评测。
  • 发现现有指标与人类判断相关性弱,尤其在长回答上表现差。
  • AURA评分显著提升与人类评价的相关性,适合研究者使用。

语音问答(AQA)是评估语音-语言模型的关键任务,但开放式回答的评估仍具挑战。现有基于NLP和音频描述的指标(如BLEU、METEOR、BERTScore)依赖表面相似性,未能考虑问题上下文、推理过程和部分正确性。本文有三项贡献:首先,提出AQEval,首个包含1万条模型生成回答并经多人人工标注正确性和相关性的基准;其次,在AQEval上全面分析现有指标,发现其与人类判断相关性较弱,尤其在长回答中;第三,提出新指标AURA Score,能更好评估开放式回答。在AQEval上,AURA显著优于所有基线,达到最先进的相关性。本文旨在揭示当前AQA评估方法的局限性,并推动更优评估体系的发展。我们已公开发布AQEval基准和AURA评分工具,以支持未来研究。

原文摘要 · Abstract (English)

Audio Question Answering (AQA) is a key task for evaluating Audio-Language Models (ALMs), yet assessing open-ended responses remains challenging. Existing metrics used for AQA such as BLEU, METEOR and BERTScore, mostly adapted from NLP and audio captioning, rely on surface similarity and fail to account for question context, reasoning, and partial correctness. To address the gap in literature, we make three contributions in this work. First, we introduce AQEval to enable systematic benchmarking of AQA metrics. It is the first benchmark of its kind, consisting of 10k model responses annotated by multiple humans for their correctness and relevance. Second, we conduct a comprehensive analysis of existing AQA metrics on AQEval, highlighting weak correlation with human judgment, especially for longer answers. Third, we propose a new metric - AURA score, to better evaluate open-ended model responses. On AQEval, AURA achieves state-of-the-art correlation with human ratings, significantly outperforming all baselines. Through this work, we aim to highlight the limitations of current AQA evaluation methods and motivate better metrics. We release both the AQEval benchmark and the AURA metric to support future research in holistic AQA evaluation.

语音问答评估指标自然语言理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。