提出细粒度音视频推理评估基准AURA,揭示模型答对题却逻辑错误的问题。
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- 设计六类认知任务,强制模型融合音视频信息进行跨模态推理。
- 92%准确率下,事实一致性与逻辑有效性得分均低于45%,暴露推理缺陷。
- 新指标AuraScore分解评估推理真实性,适合研究多模态理解的学者。
当前音视频(AV)评测聚焦最终答案正确率,忽视深层推理过程,难以区分真实理解与依赖错误逻辑或幻觉的正确答案。为此,我们提出AURA(音视频理解与推理评估),用于评估音视频大语言模型(AV-LLMs)和全模态语言模型(OLMs)的跨模态推理能力。AURA涵盖因果、音色与音高、节奏与音视频同步、不可回答性、隐含干扰、技能画像等六类挑战性认知领域,专为单模态无法解答而设计,迫使模型构建基于音频和视频的合理逻辑路径,区别于允许单模态捷径的现有数据集。为评估推理轨迹,我们提出新型指标AuraScore,分解为两方面:(i) 事实一致性——推理是否基于感知证据;(ii) 核心推理——每步推理的逻辑有效性。对最先进模型的评估显示关键推理差距:尽管准确率高达92%(部分任务),但事实一致性和核心推理得分均低于45%。这一差异表明模型常通过错误逻辑获得正确答案,凸显本基准的必要性,并为更稳健的多模态评估铺平道路。
原文摘要 · Abstract (English)
Current audio-visual (AV) benchmarks focus on final answer accuracy, overlooking the underlying reasoning process. This makes it difficult to distinguish genuine comprehension from correct answers derived through flawed reasoning or hallucinations. To address this, we introduce AURA (Audio-visual Understanding and Reasoning Assessment), a benchmark for evaluating the cross-modal reasoning capabilities of Audio-Visual Large Language Models (AV-LLMs) and Omni-modal Language Models (OLMs). AURA includes questions across six challenging cognitive domains, such as causality, timbre and pitch, tempo and AV synchronization, unanswerability, implicit distractions, and skill profiling, explicitly designed to be unanswerable from a single modality. This forces models to construct a valid logical path grounded in both audio and video, setting AURA apart from AV datasets that allow uni-modal shortcuts. To assess reasoning traces, we propose a novel metric, AuraScore, which addresses the lack of robust tools for evaluating reasoning fidelity. It decomposes reasoning into two aspects: (i) Factual Consistency - whether reasoning is grounded in perceptual evidence, and (ii) Core Inference - the logical validity of each reasoning step. Evaluations of SOTA models on AURA reveal a critical reasoning gap: although models achieve high accuracy (up to 92% on some tasks), their Factual Consistency and Core Inference scores fall below 45%. This discrepancy highlights that models often arrive at correct answers through flawed logic, underscoring the need for our benchmark and paving the way for more robust multimodal evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。