用火柴棍谜题测试视觉符号组合推理能力,发现模型普遍不如人类。
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- 通过移动1~2根火柴棍修复错误算式,融合视觉感知与符号操作。
- 生成140万实例,测试集覆盖多种难度,人类准确率超90%。
- 揭示当前视觉-语言模型在纯视觉任务中表现严重不足。
我们提出 extsc{MathSticks},一个用于视觉符号组合推理(VSCR)的基准评测。每个任务给出一个错误的火柴棍算式,需通过移动一根或两根火柴棍,在严格守恒规则下修正。该基准涵盖文本引导和纯视觉两种设置,系统性地覆盖数字规模、移动复杂度、解的多样性及运算符变化,共生成140万实例,并构建了精心筛选的测试集。对14个视觉-语言模型的评估显示:闭源模型仅在简单情况下有效,开源模型在纯视觉场景中表现失败,而人类准确率超过90%。这些结果确立了 extsc{MathSticks} 作为推动跨模态组合推理研究的严谨测试平台。代码与数据集已公开于 https://github.com/Yuheng2000/MathSticks。
原文摘要 · Abstract (English)
We introduce \textsc{MathSticks}, a benchmark for Visual Symbolic Compositional Reasoning (VSCR), which unifies visual perception, symbolic manipulation, and arithmetic consistency. Each task presents an incorrect matchstick equation that must be corrected by moving one or two sticks under strict conservation rules. The benchmark includes both text-guided and purely visual settings, systematically covering digit scale, move complexity, solution multiplicity, and operator variation, with 1.4M generated instances and a curated test set. Evaluations of 14 vision--language models reveal substantial limitations: closed-source models succeed only on simple cases, open-source models fail in the visual regime, while humans exceed 90\% accuracy. These findings establish \textsc{MathSticks} as a rigorous testbed for advancing compositional reasoning across vision and symbols. Our code and dataset are publicly available at https://github.com/Yuheng2000/MathSticks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。