测试GPT-4在3D旋转理解上的表现,发现需额外信息才能准确推理。
AI's Spatial Intelligence: Evaluating AI's Understanding of Spatial Transformations in PSVT:R and Augmented Reality
- 用修订版空间可视化测试和增强现实场景评估GPT-4的3D旋转理解能力。
- 加入坐标轴和文字/数学描述后,GPT-4准确率显著提升。
- 适合教育、制造等领域中需要空间推理支持的AI辅助系统设计。
空间智能在建筑、工程、科学、医学等STEM领域至关重要。理解三维(3D)空间旋转需借助语言描述或视觉/交互示例来展示物体方向变化。近期研究表明,具备语言与视觉能力的人工智能在空间推理方面仍存在局限。本文研究了生成式AI通过图像与语言处理理解物体旋转的能力。基于修订版普渡大学空间可视化测试:旋转可视化(Revised PSVT:R),评估GPT-4模型在理解旋转过程中的表现。进一步在测试中引入坐标系轴,观察其性能变化。同时考察了GPT-4在增强现实(AR)场景中对3D旋转的理解,发现添加描述旋转过程的文字说明或旋转矩阵等数学表示后,其理解准确率明显提高。结果表明,尽管当前主流生成式模型如GPT-4尚缺乏对空间旋转过程的内在理解,但通过外部补充信息(如AR可视化与数学表达)可显著增强其推理能力。结合AI的空间智能潜力与AR的交互可视化优势,有望为学生空间学习提供更有效指导,助力装配、制造等实际操作中的空间认知理解。
原文摘要 · Abstract (English)
Spatial intelligence is important in Architecture, Construction, Science, Technology, Engineering, and Mathematics (STEM), and Medicine. Understanding three-dimensional (3D) spatial rotations can involve verbal descriptions and visual or interactive examples, illustrating how objects change orientation in 3D space. Recent studies show Artificial Intelligence (AI) with language and vision capabilities still face limitations in spatial reasoning. In this paper, we have studied generative AI's spatial capabilities of understanding rotations of objects utilizing its image and language processing features. We examined the spatial intelligence of the GPT-4 model with vision in understanding spatial rotation process with diagrams based on the Revised Purdue Spatial Visualization Test: Visualization of Rotations (Revised PSVT:R). Next, we incorporated a layer of coordinate system axes on Revised PSVT:R to study the variations in GPT-4's performance. We also examined GPT-4's understanding of 3D rotations in Augmented Reality (AR) scenes that visualize spatial rotations of an object in 3D space and observed increased accuracy of GPT-4's understanding of the rotations by adding supplementary textual information depicting the rotation process or mathematical representations of the rotation (e.g., matrices). The results indicate that while GPT-4 as a major current Generative AI model lacks the understanding of a spatial rotation process, it has the potential to understand the rotation process with additional information that can be provided by methods such as AR. By combining the potentials in spatial intelligence of AI with AR's interactive visualization abilities, we expect to offer enhanced guidance for students' spatial learning activities. Such spatial guidance can benefit understanding spatial transformations and additionally support processes like assembly, fabrication, and manufacturing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。