评测GPT-4o在计算机图形学题中的视觉与几何推理能力,发现其潜力大但仍有明显局限。
An Eye for an AI: Evaluating GPT-4o's Visual Perception Skills and Geometric Reasoning Skills Using Computer Graphics Questions
- 构建两个包含视觉感知与几何推理的CG题目数据集
- GPT-4o在独立处理图像信息时表现良好,但结果准确率和质量仍不足
- 提出适用于计算机图形学教学的GenAI使用建议,助力课堂互动
计算机图形学(CG)是计算机科学中的热门领域,但因其涉及数学、编程、几何推理和创造力等多种技能,常令学生感到困难。近年来,研究者尝试利用生成式人工智能(GenAI)提升教学效果,但多数研究聚焦于初级编程课程。此前一项针对纯文本大语言模型GPT-4在CG问题上的评估显示其表现不佳,且高度依赖用户对图像内容的详细描述,需大量人工干预才能获得合理结果。目前尚无研究考察大型多模态模型(LMM)在解决CG问题上的能力及其教学应用潜力。本研究构建了两套涵盖不同视觉感知与几何推理难度的CG问题数据集,评估当前最先进的多模态模型GPT-4o的表现。结果表明,尽管GPT-4o在自主处理视觉信息方面展现出显著潜力,但在准确性与输出质量上仍存在明显缺陷。为此,我们提出了若干面向计算机图形学教育者的新型应用策略,以在现有局限下有效整合GenAI,促进课堂教学中的学习参与度。
原文摘要 · Abstract (English)
CG (Computer Graphics) is a popular field of CS (Computer Science), but many students find this topic difficult due to it requiring a large number of skills, such as mathematics, programming, geometric reasoning, and creativity. Over the past few years, researchers have investigated ways to harness the power of GenAI (Generative Artificial Intelligence) to improve teaching. In CS, much of the research has focused on introductory computing. A recent study evaluating the performance of an LLM (Large Language Model), GPT-4 (text-only), on CG questions, indicated poor performance and reliance on detailed descriptions of image content, which often required considerable insight from the user to return reasonable results. So far, no studies have investigated the abilities of LMMs (Large Multimodal Models), or multimodal LLMs, to solve CG questions and how these abilities can be used to improve teaching. In this study, we construct two datasets of CG questions requiring varying degrees of visual perception skills and geometric reasoning skills, and evaluate the current state-of-the-art LMM, GPT-4o, on the two datasets. We find that although GPT-4o exhibits great potential in solving questions with visual information independently, major limitations still exist to the accuracy and quality of the generated results. We propose several novel approaches for CG educators to incorporate GenAI into CG teaching despite these limitations. We hope that our guidelines further encourage learning and engagement in CG classrooms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。