arXiv:2608.14741cs.CVcs.AI2026-08

测试多模态模型组合3D空间推理能力,用积木块拼图挑战。

PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models

论文配图:PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models
图 1 · 摘自论文原文
  • 用程序生成120个积木拼合难题,验证视觉与组合推理能力。
  • 最强模型仅达50%准确率,远低于人类水平,说明当前模型仍有不足。
  • 适合评估多模态模型在复杂空间关系理解上的表现,尤其关注3D推理。

我们提出PolyComp,一个程序生成并验证的基准,用于检验视觉识别与组合空间推理能力。每个问题要求模型从四个选项中选出能拼合成目标立体的两个多立方体组件。该基准包含120个问题,覆盖四个几何类别,每道题有单图或多图三种呈现方式。随机猜测基线为25%。在三种呈现方式下(每模型共360题),GPT-5.6 Sol(全力)达到50.0%准确率(95%问题簇置信区间43.3%-56.7%),平均成本$0.951;Claude Fable 5(全力)达39.4%(33.1%-46.1%),成本$0.701;Gemini 3.1 Pro Preview(高思考层级)仅27.5%(22.8%-32.5%),接近随机水平,成本$0.350。不同几何类别的准确率差异大于不同呈现方式间的差异。我们提供问题开发与评估流程、成本与令牌统计,并公开全部120个问题。

原文摘要 · Abstract (English)

We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that can be combined to form a target solid. The benchmark contains 120 problems across four geometry families, and each problem has three different presentation formats using either a single image or multiple images. The random guessing baseline is 25%. Across the three presentations (360 presented problems per model), GPT-5.6 Sol with max effort attains 50.0% accuracy (95% problem-cluster CI 43.3-56.7%) at a mean cost of \$0.951 per presented problem, Claude Fable 5 with max effort attains 39.4% (33.1-46.1%) at \$0.701, and Gemini 3.1 Pro Preview with thinking level high attains 27.5% (22.8-32.5%), near the 25% random guessing baseline, at \$0.350. The observed accuracy spread across geometry families is larger than across presentation formats. We present a problem development and evaluation protocol, cost and token accounting, and release the 120 problems.

3D推理多模态积木拼合基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。