评测大模型在多路径探索上的推理能力,发现现有模型难兼顾深度与广度。
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
- 构建1200+跨领域多模态案例集,设计细粒度树状思维评估协议。
- 30+主流多模态大模型测试中,普遍在宽广探索与深度链式推理间表现不足。
- 揭示失败模式,为打造更全面推理能力的模型提供方向,适合研究者参考。
本文提出一个全面的多模态基准,专门评估多模态大模型(MLLMs)的推理宽度,这一维度与常被研究的推理深度形成互补。推理深度衡量模型进行长链、强关联的顺序推理能力;而推理宽度则关注模型在并行路径中进行广泛试错或多重约束优化的能力,需系统遍历多种可能路径,应用不同约束剪枝无效分支,并识别有效解题路径以实现高效迭代或回溯。为此,我们精心构建了超过1200个高质量多模态案例,覆盖异构领域,并提出一种细粒度的树状思维(Tree-of-Thought)评估协议,可联合量化推理宽度与深度。我们在多个难度层级、问题类型和所需技能下,对12个主要模型家族(共30余个先进模型)进行了评估。结果表明,尽管当前模型在通用或常识性视觉问答任务中表现良好,但在结合深序列思维链与广范围探索以实现真正基于洞察的推理方面仍存在明显短板。最后,我们分析了典型失败模式,为构建既更深又更广的推理能力提供可能方向。
原文摘要 · Abstract (English)
In this paper, we present a holistic multimodal benchmark that evaluates the reasoning capabilities of MLLMs with an explicit focus on reasoning width, a complementary dimension to the more commonly studied reasoning depth. Specifically, reasoning depth measures the model's ability to carry out long-chain, sequential reasoning in which each step is tightly and rigorously linked to the next. Reasoning width tends to focus more on the model's capacity for broad trial-and-error search or multi-constrained optimization: it must systematically traverse many possible and parallelized reasoning paths, apply diverse constraints to prune unpromising branches, and identify valid solution routes for efficient iteration or backtracking. To achieve it, we carefully curate 1200+ high-quality multimodal cases spanning heterogeneous domains, and propose a fine-grained tree-of-thought evaluation protocol that jointly quantifies reasoning width and depth. We evaluate 12 major model families (over 30 advanced MLLMs) across difficulty tiers, question types, and required skills. Results show that while current models exhibit strong performance on general or common-sense VQA tasks, they still struggle to combine deep sequential thought chains with wide exploratory search to perform genuine insight-based reasoning. Finally, we analyze characteristic failure modes to provide possible directions for building MLLMs that reason not only deeper but also wider.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。