提出新基准,评估模型如何智能选择视觉或文本推理
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
- 基于模型能力动态判断任务难度,评估模式选择合理性
- 发现模型越强越会自适应选择,但准确率不必然提升
- 适合研究多模态推理策略与模型决策过程的学者
自适应多模态推理已成为视觉语言模型的新前沿,旨在动态切换工具增强的视觉推理与文本推理,以提升效果与效率。然而,现有评估依赖静态难度标签和简单指标,无法捕捉难度随模型能力变化的动态特性,混淆了自适应选择与整体性能,并忽视细粒度过程分析。本文提出 AdaptMMBench,一个覆盖现实世界、OCR、GUI、知识与数学五大领域的综合性基准,包含直接感知与复杂推理任务。该基准采用马修斯相关系数(MCC)评估不同推理模式的选择合理性,通过动态识别任务难度与模型能力边界,分离出元认知能力。此外,支持多维度过程评估,涵盖关键步骤覆盖率、工具有效性与计算效率。评估显示:自适应模式选择随模型容量提升,但显著脱离最终准确率;关键步骤覆盖率与性能一致,但工具有效性在不同架构间差异极大。
原文摘要 · Abstract (English)
Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text reasoning to enhance both effectiveness and efficiency. However, existing evaluations rely on static difficulty labels and simplistic metrics, which fail to capture the dynamic nature of difficulty relative to varying model capacities. Consequently, they obscure the distinction between adaptive mode selection and general performance while neglecting fine-grained process analyses. In this paper, we propose AdaptMMBench, a comprehensive benchmark for adaptive multimodal reasoning across five domains: real-world, OCR, GUI, knowledge, and math, encompassing both direct perception and complex reasoning tasks. AdaptMMBench utilizes a Matthews Correlation Coefficient (MCC) metric to evaluate the selection rationality of different reasoning modes, isolating this meta-cognition ability by dynamically identifying task difficulties based on models' capability boundaries. Moreover, AdaptMMBench facilitates multi-dimensional process evaluation across key step coverage, tool effectiveness, and computational efficiency. Our evaluation reveals that while adaptive mode selection scales with model capacity, it notably decouples from final accuracy. Conversely, key step coverage aligns with performance, though tool effectiveness remains highly inconsistent across model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。