测试大模型画图辅助解题能力,发现高正确率背后是推理幻觉。
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
- 设计三阶段评估框架,检验画图和基于图推理的能力。
- 实测主流模型在画图环节失败率超80%,但答案准确率仍达70%。
- 适合关注多模态模型真实推理能力的研究者与开发者。
高级人工智能的标志是从被动视觉感知向战略性修改视觉信息以促进复杂推理的转变。然而,当前的大规模多模态模型(LMMs)在这一能力上严重不足。该缺陷常被仅关注最终答案准确率的评估指标所掩盖,造成虚假的智能表象。我们以几何问题求解为精确工具,通过需构建视觉辅助的任务来揭示此问题。为此,我们提出 extbf{VisAidMath} 基准测试及创新的三层次漏斗评估框架。该框架超越简单准确率(ACCU),深入检验视觉辅助生成有效性(PVA)和后续推理步骤合理性(SPRS)。对包括 Doubao-Seed-1.6 和 o4 等前沿模型的广泛实验显示,存在深刻的“推理幻觉”:表面准确率高,实则模型无法生成有效视觉辅助或基于其进行合理推理。研究揭示现代 LMMs 在视觉感知与逻辑推导间存在根本断裂。评估平台已上线 CodaBench,可供公开测试。主页:https://nlp2ct.github.io/VisAidMathHomepage/ 评测页:https://www.codabench.org/competitions/7634/
原文摘要 · Abstract (English)
A hallmark of advanced artificial intelligence is the capacity to progress from passive visual perception to the strategic modification of visual information to facilitate complex reasoning. This advanced capability, however, remains critically underdeveloped in current Large Multi-modal Models (LMMs). The deficiency is often masked by evaluation metrics that prioritize final-answer accuracy, creating an illusion of competence where genuine reasoning is absent. Using the domain of geometric problem-solving as a precise instrument, we probe this issue through tasks that require constructing visual aids. To this end, we introduce \textbf{VisAidMath}, a challenging benchmark, and our novel Three-Layered Funnel Evaluation Framework. This framework moves beyond simple accuracy (ACCU) to scrutinize the generation of valid visual aids (PVA) and the soundness of subsequent reasoning steps (SPRS). Our extensive experiments on state-of-the-art models, including Doubao-Seed-1.6 and o4, reveal a profound ``Reasoning Illusion''. We observe that high surface-level accuracy conceals a catastrophic failure in the models' ability to produce valid visual aids or to reason from them. Our findings expose a fundamental schism between visual perception and logical deduction in modern LMMs. We host an evaluation platform at CodaBench for testing publicly. Homepage: https://nlp2ct.github.io/VisAidMathHomepage/ Evaluation: https://www.codabench.org/competitions/7634/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。