测试视觉模型对三维几何形状的理解能力,发现现有模型普遍存在认知缺陷。
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- 构建包含仿真与真实多面体的综合性基准测试集
- 主流模型在重建柏拉图固体时准确率低,识别对称性能力弱
- 适合研究几何理解、鲁棒3D表示学习的学者使用
现代单目三维重建方法和视觉语言模型(VLMs)在标准基准上表现优异,但近期研究质疑其对几何属性的真实理解能力。我们提出GIQ,一个专门用于评估视觉与视觉-语言基础模型几何推理能力的综合基准。GIQ包含多样化的合成与真实图像及对应3D网格,涵盖从柏拉图、阿基米德、约翰逊、卡塔兰立体到星形化与复合结构的各类多面体,覆盖不同复杂度与对称性。通过系统实验,包括单目三维重建、三维对称性检测、心理旋转测试与零样本形状分类任务,我们揭示当前模型存在显著不足:在大规模3D数据集上训练的先进重建算法仍无法准确重建基本柏拉图固体;尽管基础模型可通过线性与非线性探针捕捉特定三维对称元素,但在需要精细几何区分的任务中表现差;此外,ChatGPT、Gemini、Claud等先进视觉语言助手在解释复杂多面体的基本几何属性如面结构、凸性与复合结构时准确率极低。GIQ已公开于toomanymatts.github.io/giq-benchmark/,为衡量几何智能关键差距提供结构化平台,助力未来鲁棒、几何感知表征学习的发展。
原文摘要 · Abstract (English)
Modern monocular 3D reconstruction methods and vision-language models (VLMs) demonstrate impressive results on standard benchmarks, yet recent works cast doubt on their true understanding of geometric properties. We introduce GOQ, a comprehensive benchmark specifically designed to evaluate the geometric reasoning capabilities of vision and vision-language foundation models. GIQ comprises synthetic and real-world images and corresponding 3D meshes of diverse polyhedra covering varying levels of complexity and symmetry, from Platonic, Archimedean, Johnson, and Catalan solids to stellations and compound shapes. Through systematic experiments involving monocular 3D reconstruction, 3D symmetry detection, mental rotation tests, and zero-shot shape classification tasks, we reveal significant shortcomings in current models. State-of-the-art reconstruction algorithms trained on extensive 3D datasets struggle to reconstruct even basic geometric Platonic solids accurately. Next, although foundation models may be shown via linear and non-linear probing to capture specific 3D symmetry elements, they falter significantly in tasks requiring detailed geometric differentiation, such as mental rotation. Moreover, advanced vision-language assistants such as ChatGPT, Gemini and Claud exhibit remarkably low accuracy in interpreting basic shape properties such as face geometry, convexity, and compound structures of complex polyhedra. GIQ is publicly available at toomanymatts.github.io/giq-benchmark/, providing a structured platform to benchmark critical gaps in geometric intelligence and facilitate future progress in robust, geometry-aware representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。