用画钟测试发现生成式AI存在认知缺陷,虽能画出时钟外形但难准确定时。
Evidence of Cognitive Deficits andDevelopmental Advances in Generative AI: A Clock Drawing Test Analysis
- 通过画钟测试评估生成式AI的视觉空间与规划能力。
- 仅GPT-4 Turbo和Gemini Pro 1.5正确画出时间,其他模型普遍出错。
- 适合关注AI认知局限与人机认知对比的研究者参考。
生成式AI的快速发展引发了对其认知能力的兴趣,尤其在语言理解与代码生成等任务中的表现。本研究探讨了多个近期生成式AI模型在钟表绘制测试(Clock Drawing Test, CDT)上的表现,该测试是评估视觉空间规划与组织能力的神经心理学工具。尽管模型能绘制出类似时钟的图形,但在时间表示上存在显著缺陷,错误模式类似于轻度至重度认知障碍(Wechsler, 2009)。常见错误包括数字顺序混乱、时间显示错误及添加无关元素,尽管时钟结构本身绘制准确。只有GPT-4 Turbo和Gemini Pro 1.5成功绘制出正确时间,得分相当于健康个体(4/4)。后续钟表读数测试显示,仅有Sonnet 3.5成功完成,表明绘图缺陷源于对数字概念的理解困难。这些结果可能反映视觉空间理解、工作记忆或计算能力的不足,凸显其在已学知识上的优势与推理能力的薄弱。比较人类与机器的表现对理解人工智能的认知能力至关重要,并有助于引导其向类人认知功能发展。
原文摘要 · Abstract (English)
Generative AI's rapid advancement sparks interest in its cognitive abilities, especially given its capacity for tasks like language understanding and code generation. This study explores how several recent GenAI models perform on the Clock Drawing Test (CDT), a neuropsychological assessment of visuospatial planning and organization. While models create clock-like drawings, they struggle with accurate time representation, showing deficits similar to mild-severe cognitive impairment (Wechsler, 2009). Errors include numerical sequencing issues, incorrect clock times, and irrelevant additions, despite accurate rendering of clock features. Only GPT 4 Turbo and Gemini Pro 1.5 produced the correct time, scoring like healthy individuals (4/4). A follow-up clock-reading test revealed only Sonnet 3.5 succeeded, suggesting drawing deficits stem from difficulty with numerical concepts. These findings may reflect weaknesses in visual-spatial understanding, working memory, or calculation, highlighting strengths in learned knowledge but weaknesses in reasoning. Comparing human and machine performance is crucial for understanding AI's cognitive capabilities and guiding development toward human-like cognitive functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。