评测大模型在跨模态任务整合与复杂指令理解上的短板
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
- 构建1005个真实世界多模态任务,评估模型融合视觉语言能力的水平
- 顶尖模型(Gemini 2.5 Pro)仅达44%准确率,远低于实际应用要求
- 首次系统评测复杂文本与视觉指令的对齐能力,助力未来模型优化
大型多模态模型(LMMs)在视觉-语言(VL)任务中展现出强大的通用性潜力,但在需要整合多种视觉语言能力或理解复杂文本/视觉指令的任务中表现不佳。为深入探究这一差距及其成因,我们提出MOAT——一个包含1005个复杂真实世界视觉问题的基准测试。这些任务对人类而言简单明了,但对现有模型极具挑战性。任务要求模型综合运用阅读文本、计数、空间关系理解、指令对齐等9类视觉语言能力,该分类体系使我们能细致分析模型优劣势。此外,MOAT是首个专门评估复杂文本与视觉指令对齐能力的基准,这对实际应用至关重要。我们评估了17个开源及专有模型,发现表现最佳者(Gemini 2.5 Pro)准确率仅为44%,远未达到实用标准。通过分析结果趋势,我们揭示了以文本为中心推理的局限性、关键能力瓶颈以及拼贴策略可能带来的负面影响,为未来模型发展提供方向。代码与数据已公开于https://cambrian-yzt.github.io/MOAT/。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by their poor performance in tasks that require a combination of VL capabilities, as well as in tasks that involve the grounding of complex text or visual instructions. To thoroughly investigate this gap and its underlying causes, we propose MOAT, a diverse benchmark with 1005 complex real-world vision questions that are straightforward for humans but challenging for LMMs. Specifically, the tasks in MOAT require LMMs to engage in generalist problem solving by integrating VL capabilities such as reading text, counting, understanding spatial relations, grounding textual and visual instructions, etc. All these abilities fit into a taxonomy proposed by us that contains 9 VL capabilities, enabling MOAT to provide a fine-grained view of LMMs' strengths and weaknesses. Besides, MOAT is the first benchmark to explicitly evaluate LMMs' ability to ground complex text and visual instructions, which is essential for many real-world applications. We evaluated 17 proprietary and open source LMMs, finding that the best performing LMM (Gemini 2.5 Pro) achieved only 44% accuracy, far below what would be acceptable in real-world applications. To guide future model development, we analyze common trends in our results and discuss the underlying causes of poor performance, focusing on the impact of text-centric reasoning, which VL capabilities form bottlenecks in complex tasks, and the potential harmful effects of tiling. Code and data are available at https://cambrian-yzt.github.io/MOAT/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。