用拼图任务评估多模态大模型的空间组合推理能力。
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
- 设计可机器验证的几何拼图表达式,支持多重正确解。
- 发现模型常忽略几何约束,导致拼图变形。
- 适合研究空间推理与视觉生成的学者使用。
多模态大语言模型在视觉识别和语义理解方面取得了显著进展,但在几何约束下的精确组合空间推理仍缺乏深入探索。现有基准主要评估粗略的空间关系,很少支持严格的几何验证或构造任务中的多重有效解。为解决这些问题,我们提出TangramPuzzle,一个包含668个已验证配置和1,336个实例的基准,用于评估在几何约束下的组合空间推理能力。我们提出了拼图构造表达式(TCE),以机器可验证的几何规范编码拼图配置,并设计了多解感知验证器,分别评估几何正确性和轮廓匹配度。TangramPuzzle涵盖两个互补任务:部件到整体的轮廓预测,以及整体到部件的端到端拼装生成。我们在先进开源和专有模型上进行了广泛评估,揭示了一个有趣现象:多模态大模型倾向于优先匹配目标轮廓,而忽视几何约束,导致拼图块出现扭曲或变形。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition and semantic understanding, yet precise compositional spatial reasoning under geometric constraints remains underexplored. Existing benchmarks mainly assess coarse spatial relations and rarely support rigorous geometric verification or multiple valid solutions in constructive tasks. To address these limitations, we introduce TangramPuzzle, a benchmark comprising 668 validated configurations and 1,336 instances for evaluating compositional spatial reasoning under geometric constraints. We formulate the Tangram Construction Expression (TCE) to encode tangram configurations with machine-verifiable geometric specifications, together with a multi-solution-aware verifier that separately evaluates geometric validity and silhouette fidelity. TangramPuzzle covers two complementary evaluation tasks: part-to-whole Outline Prediction and whole-to-part End-to-End Assembly Generation. We conduct extensive evaluation experiments on advanced open-source and proprietary models, revealing an interesting insight: MLLMs tend to prioritize matching the target silhouette while neglecting geometric constraints, leading to distortions or deformations of the pieces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。