arXiv:2503.09321cs.CVcs.AI2025-03NeurIPS被引 5

DAVE基准测试让音视频模型问题一目了然

DAVE: Diagnostic benchmark for Audio Visual Evaluation

  • 设计双模态必要任务,确保音视频都必须用上
  • 拆解评估为原子子任务,定位模型具体短板
  • 适合想提升音视频对齐能力的研究者

音频视觉理解是快速发展的领域,旨在融合和解释听觉与视觉信息。尽管多模态学习取得进展,现有基准常存在强视觉偏差——答案可仅从视觉数据推断——且仅提供混合评分,混淆了多种错误来源。这使得难以判断模型在视觉理解、音频解析或音视频对齐上的真实表现。本文提出DAVE:诊断性音视频评估基准,一个新设计的基准数据集,可在受控环境下系统评估音视频模型。DAVE通过(i)确保两模态均需正确作答,(ii)将评估拆分为原子子类别,缓解上述局限。对前沿模型的详细分析揭示了特定失败模式,并提供针对性改进见解。通过提供标准化诊断框架,我们旨在推动音视频模型更稳健发展。

原文摘要 · Abstract (English)

Audio-visual understanding is a rapidly evolving field that seeks to integrate and interpret information from both auditory and visual modalities. Despite recent advances in multi-modal learning, existing benchmarks often suffer from strong visual bias -- when answers can be inferred from visual data alone -- and provide only aggregate scores that conflate multiple sources of error. This makes it difficult to determine whether models struggle with visual understanding, audio interpretation, or audio-visual alignment. In this work, we introduce DAVE: Diagnostic Audio Visual Evaluation, a novel benchmark dataset designed to systematically evaluate audio-visual models across controlled settings. DAVE alleviates existing limitations by (i) ensuring both modalities are necessary to answer correctly and (ii) decoupling evaluation into atomic subcategories. Our detailed analysis of state-of-the-art models reveals specific failure modes and provides targeted insights for improvement. By offering this standardized diagnostic framework, we aim to facilitate more robust development of audio-visual models. Dataset: https://huggingface.co/datasets/gorjanradevski/dave Code: https://github.com/gorjanradevski/dave

音视频理解多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。