arXiv:2607.27670cs.CVcs.AI2026-07

用拼图测试视觉-几何推理,发现主流模型缺乏几何理解能力。

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

论文配图:JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
图 1 · 摘自论文原文
  • 设计带凹凸结构的拼图块,结合视觉与几何约束确保答案唯一。
  • 零样本下仅1个模型在4×4拼图上超随机水平,大网格性能暴跌。
  • 揭示当前模型在复杂几何推理上的瓶颈,适合关注多模态认知的学者。

拼图求解需同时处理视觉内容与几何约束,但现有基准使用矩形切割,导致纹理重复区域存在模糊真值。我们提出 extit{ ous{}},采用带凹槽与空缺的互锁拼图块,几何约束提供强局部兼容性要求,结合视觉信息可获得无歧义真值。在95,000个实例、四种网格密度(4×4至16×16)上,发现零样本视觉-语言模型(VLMs)普遍缺乏几何推理能力:五种前沿模型中仅一个(GPT-5.5)在4×4拼图上超过随机基线,其余均处于随机水平。尽管监督微调可在4×4达到97%以上准确率,但所有模型在更大网格上崩溃:GPT-5.5从70%降至近随机水平(8×8),即使微调模型在12×12上也低于5%。这一‘缩放悬崖’表明当前架构无法随拼图块数增加维持一致的约束满足。 extit{ ous{}}确立了可扩展几何推理作为视觉-语言模型的一项开放挑战。

原文摘要 · Abstract (English)

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.

视觉推理拼图任务几何约束多模态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。