现有多模态大模型评估缺乏跨模态整合能力测试。
What We are Missing in Multimodal LLM Evaluation?

- 分析现有评测方法,发现跨模态整合能力评估缺失
- 指出时间-空间连贯性等四大关键评估缺口
- 适合关注多模态模型真实智能水平的研究者
多模态大语言模型(MLLMs)可处理文本、图像、音频、视频等多种输入并生成文本响应。尽管其能力迅速提升,但评估体系未能同步发展。现有评测基准大多局限于单一任务,难以反映模型在多模态信息融合方面的表现。本文分析当前评估手段,梳理现有基准分类体系,揭示了时间-空间连贯性、物理世界理解、多模态一致性及选择性注意力等关键评估缺口。弥补这些不足对于衡量多模态智能的真实进展、揭示模型能力边界至关重要。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced rapidly, evaluation of such models has not kept pace. Most existing evaluation benchmarks are limited to isolated tasks and reveal little about whether a model integrates information across modalities. We examine current means for evaluating MLLMs and review the existing benchmark taxonomy to identify gaps, including temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention. Addressing these gaps is essential for measuring real progress in multimodal intelligence and exposing capability boundaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。