构建音视频智能评估基准,诊断多模态大模型短板
AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

- 设计三阶段跨模态任务,实现音视频联合理解评估
- 发现当前多模态模型在音视频推理上存在显著不足
- 适合研究音视频感知与通用性建模的学者参考
近期多模态大语言模型在视觉、音频与语言融合方面取得进展,但其音视频智能(AVI)仍缺乏系统性评估。本文提出AVI-Bench,一个受认知启发的基准,通过跨模态任务在感知、理解、推理三个阶段评估模型的音视频联合解释能力,实现细粒度的能力诊断与失效分析。为检验模型在陌生场景下的鲁棒性,进一步提出AVI-Bench-PriSe,使用低语义刺激测试原始音视频感知能力,评估模型对训练分布外数据的泛化性能。对开源与闭源模型的广泛实验揭示了当前多模态模型在音视频推理方面的显著局限。基于结果,提出四级音视频智能分类体系。总体而言,AVI-Bench提供了一个原则性的评估框架,可指导更具鲁棒性和泛化性的音视频智能发展。
原文摘要 · Abstract (English)
Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a cognitively inspired benchmark that evaluates Omni-MLLMs across three stages, perception, understanding, and reasoning, through cross-modal tasks requiring joint audio-visual interpretation. This design enables fine-grained diagnosis of model capabilities and failure modes. To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models' primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. Extensive experiments on both open-source and closed-source models reveal substantial limitations in current Omni-MLLMs. Based on these findings, we present a four-level AVI taxonomy. Overall, AVI-Bench provides a principled evaluation framework to guide the development of more robust and generalizable AVI. Project website: https://fudancvl.github.io/AVI-Bench/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。