构建视觉网格推理谜题基准,评估大模型逻辑解题能力。
VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
- 设计20个不同难度的视觉网格谜题,覆盖多类逻辑规则。
- 发现顶尖大模型在谜题上仍表现不佳,存在本质缺陷。
- 提出两种微调策略,提升特定谜题表现但泛化有限。
大型视觉语言模型(LVLMs)在需要精确感知、规则理解与逻辑推理的谜题任务中表现薄弱。现有评测基准多针对预训练模型,缺乏对推理能力的专注,且未建立系统性评估框架。为此,本文提出VGRP-Bench,一个包含20种多样化视觉网格推理谜题的基准,覆盖多级难度。实验不仅涵盖GPT-4o等主流聊天型LVLM,也包括Gemini-Thinking等推理优化型模型。结果表明,即使最先进的模型在该任务上仍表现不足,暴露出其在结构化推理上的根本局限。通过系统性分析,我们识别出线索数量、网格尺寸和规则复杂度是影响性能的关键因素。此外,我们探索了两种后训练微调策略:基于解法的监督微调(S-SFT)和基于合成推理过程的微调(R-SFT),二者虽显著提升对已训练谜题的准确率,但在未见谜题上泛化能力有限。项目页面:https://yufan-ren.com/subpage/VGRP-Bench/。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) struggle with puzzles, which require precise perception, rule comprehension, and logical reasoning. Assessing and enhancing their performance in this domain is crucial, as it reflects their ability to engage in structured reasoning - an essential skill for real-world problem-solving. However, existing benchmarks primarily evaluate pre-trained models without additional training or fine-tuning, often lack a dedicated focus on reasoning, and fail to establish a systematic evaluation framework. To address these limitations, we introduce VGRP-Bench, a Visual Grid Reasoning Puzzle Benchmark featuring 20 diverse puzzles. VGRP-Bench spans multiple difficulty levels, and includes extensive experiments not only on existing chat LVLMs (e.g., GPT-4o), but also on reasoning LVLMs (e.g., Gemini-Thinking). Our results reveal that even the state-of-the-art LVLMs struggle with these puzzles, highlighting fundamental limitations in their puzzle-solving capabilities. Most importantly, through systematic experiments, we identify and analyze key factors influencing LVLMs' puzzle-solving performance, including the number of clues, grid size, and rule complexity. Furthermore, we explore two Supervised Fine-Tuning (SFT) strategies that can be used in post-training: SFT on solutions (S-SFT) and SFT on synthetic reasoning processes (R-SFT). While both methods significantly improve performance on trained puzzles, they exhibit limited generalization to unseen ones. We will release VGRP-Bench to facilitate further research on LVLMs for complex, real-world problem-solving. Project page: https://yufan-ren.com/subpage/VGRP-Bench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。