arXiv:2505.20672cs.AI2025-05被引 1

用GIF图像构建类比数据集,提升AI的抽象推理能力

GIFARC: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate AI Reasoning

  • 用LLM和VLM从GIF中合成带类比关系的ARC任务
  • 引入类比引导后,模型准确率显著提升,更接近人类解题方式
  • 适合研究认知启发式推理与人机协同智能的学者

抽象与推理基准(ARC)对通用人工智能能力提出了严峻挑战,要求模型仅凭少量示例推断抽象模式。尽管深度学习取得进展,当前最佳模型在2024年ARC竞赛中的准确率仍仅为40%-55%,远低于人类水平。本文提出一种基于类比的新型合成数据集GIFARC,利用大语言模型(LLMs)和视觉-语言模型(VLMs),从包含类比关系的GIF图像中生成全新ARC风格任务,并为每项任务提供真实类比标注,明确建立视觉变换与日常概念之间的映射。通过在任务中嵌入强健的人类直观类比,GIFARC引导AI代理在进行暴力模式搜索前先进行类比推理,从而有效降低问题复杂度,形成更简洁、可解释的解决方案。实证表明,使用GIFARC引导的LLM其解题策略明显趋近于人类的类比思维模式。

原文摘要 · Abstract (English)

The Abstraction and Reasoning Corpus (ARC) poses a stringent test of general AI capabilities, requiring solvers to infer abstract patterns from only a handful of examples. Despite substantial progress in deep learning, state-of-the-art models still achieve accuracy rates of merely 40-55% on 2024 ARC Competition, indicative of a significant gap between their performance and human-level reasoning. In this work, we seek to bridge that gap by introducing an analogy-inspired ARC dataset, GIFARC. Leveraging large language models (LLMs) and vision-language models (VLMs), we synthesize new ARC-style tasks from a variety of GIF images that include analogies. Each new task is paired with ground-truth analogy, providing an explicit mapping between visual transformations and everyday concepts. By embedding robust human-intuitive analogies into ARC-style tasks, GIFARC guides AI agents to evaluate the task analogically before engaging in brute-force pattern search, thus efficiently reducing problem complexity and build a more concise and human-understandable solution. We empirically validate that guiding LLM with analogic approach with GIFARC affects task-solving approaches of LLMs to align with analogic approach of human.

类比推理视觉理解数据合成AI认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。