arXiv:2410.19546cs.AIcs.LG2024-10ICML被引 16

测试AI看图推理能力,发现大模型仍难解经典视觉谜题

Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?

  • 用经典Bongard视觉谜题检验多模态模型的抽象推理能力
  • 即使简单图案如螺旋线也常判断错误,正确率远低于人类
  • 暴露模型对基础视觉概念理解不足,难以泛化到新问题

近期出现的视觉-语言模型(如OpenAI的o1)似乎在跨模态推理中展现出先进能力。然而,其在语言引导感知与抽象推理方面的深度仍不明确,是否真能兑现承诺尚存疑问。为评估进展并揭示短板,我们引入经典的Bongard问题——一组需类人模式识别与抽象推理能力的视觉谜题。通过大规模评测,我们发现尽管部分模型偶尔能识别出区分性概念并解决个别问题,但总体表现仍不稳定。令人意外的是,对人类而言极为简单的概念(如基本螺旋线)对模型仍是挑战。当被明确要求识别真实概念时,模型依然频繁出错,表明其不仅缺乏对基础视觉概念的理解,也无法泛化至未见概念。与人类表现对比显示,机器认知与人类视觉推理间仍存在显著差距。

原文摘要 · Abstract (English)

Recently, newly developed Vision-Language Models (VLMs), such as OpenAI's o1, have emerged, seemingly demonstrating advanced reasoning capabilities across text and image modalities. However, the depth of these advances in language-guided perception and abstract reasoning remains underexplored, and it is unclear whether these models can truly live up to their ambitious promises. To assess the progress and identify shortcomings, we enter the wonderland of Bongard problems, a set of classic visual reasoning puzzles that require human-like abilities of pattern recognition and abstract reasoning. With our extensive evaluation setup, we show that while VLMs occasionally succeed in identifying discriminative concepts and solving some of the problems, they frequently falter. Surprisingly, even elementary concepts that may seem trivial to humans, such as simple spirals, pose significant challenges. Moreover, when explicitly asked to recognize ground truth concepts, they continue to falter, suggesting not only a lack of understanding of these elementary visual concepts but also an inability to generalize to unseen concepts. We compare the results of VLMs to human performance and observe that a significant gap remains between human visual reasoning capabilities and machine cognition.

视觉推理多模态模型抽象思维Bongard问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。