测试大模型用画图解题的能力,发现直接算反而更准。
VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

- 设计新基准,让模型先画图再推理
- 1168道题中,画图解法整体不如直接计算
- 适合研究多模态推理与可视化工具的局限性
多模态大模型在复杂推理上能力日益增强,但在需借助工具外化问题并基于工具输出推理时表现下降,尤其依赖视觉辅助时更为明显。这一差距至关重要,因工程与科学工作流程常依赖可视化工具进行分析、验证和决策。为此,我们提出VAMPS(视觉辅助数学问题求解基准),一个基于图表的数学评测集。VAMPS包含1,168个双语多选题,源自伊朗大学入学考试的代数与微积分题目,并扩展了经人工审核的LLM生成合成变体,所有题目均满足绘图可自然引导解题(如揭示交点、极值、渐近线等)。该基准不仅用于评测,还可诊断模型表现。它超越以往仅评估对固定图像的推理,而是测试模型能否通过构建有效图表并基于可视化结果得出答案。我们发现,尽管绘图是自然策略,但跨多种模型,直接解析解法仍优于依赖工具的视觉解法。
原文摘要 · Abstract (English)
Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a problem through a tool and then reason over the tool's output, specifically when they rely on visual aids. This gap is especially important because real engineering and scientific workflows often rely on visualization tools for analysis, validation, and decision-making. To study this discrepancy, we introduce VAMPS (Visual-Assisted Mathematical Problem Solving), a benchmark for graph-assisted mathematics. VAMPS contains 1,168 multimodal, bilingual multiple-choice question-answer pairs drawn from Iranian University Entrance Exam algebra and calculus problems and expanded with human-reviewed LLM-generated synthetic variants, all selected so that plotting provides a natural solution strategy by revealing intersections, extrema, asymptotes, etc. Designed for both benchmarking and diagnosis, VAMPS goes beyond prior multimodal benchmarks that primarily evaluate reasoning over fixed visual inputs by testing whether a model can benefit from constructing a useful graph and grounding its answer in the resulting visualization. Overall, we found that across a diverse set of models, direct analytical solving surprisingly outperforms tool-enabled visual solving, even on problems where plotting is a natural strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。