生成更丰富的抽象推理任务数据,助力通用AI能力评估
ARC-GEN: A Mimetic Procedural Benchmark Generator for the Abstraction and Reasoning Corpus
- 基于原始数据分布,程序化生成400个任务的输入输出对
- 扩展了原数据集样本量,支持更高效的任务学习验证
- 适用于算法竞赛与通用智能评测,提升基准可信度
抽象与推理语料库(ARC-AGI)是衡量人工智能向通用智能演进的重要挑战性基准。与其他评估特定任务技能或知识积累的数据集不同,ARC-AGI专注于衡量技能习得效率,而这一特性在当前最先进机器学习系统中仍显著缺失。由于每个任务仅提供少量(若干组)输入输出网格作为示范,限制了算法训练所需样本数量。本文提出ARC-GEN,一个开源的程序化生成器,旨在尽可能忠实扩展原始的ARC-AGI训练数据集。该生成器覆盖全部400个任务,且具备仿生特性——严格遵循初始版本(ARC-AGI-1)的数据分布特征。我们还探讨其在2025年谷歌代码高尔夫大赛中构建静态基准套件的应用,用于验证提交程序的正确性。
原文摘要 · Abstract (English)
The Abstraction and Reasoning Corpus remains one of the most compelling and challenging benchmarks for tracking progress toward achieving Artificial General Intelligence. In contrast to other evaluation datasets designed to assess an agent's task-specific skills or accumulated knowledge, the ARC-AGI suite is specifically targeted at measuring skill acquisition efficiency, a trait that has (so far) been lacking in even the most sophisticated machine learning systems. For algorithms that require extensive intra-task exemplars, a significant constraint imposed by ARC-AGI is the modest cardinality of its demonstration set, comprising a small number of $\langle$ input, output $\rangle$ grids per task specifying the corresponding transformation. To embellish the space of viable sample pairs, this paper introduces ARC-GEN, an open-source procedural generator aimed at extending the original ARC-AGI training dataset as faithfully as possible. Unlike prior efforts, our generator is both exhaustive (covering all four-hundred tasks) and mimetic (more closely honoring the distributional properties and characteristics embodied in the initial ARC-AGI-1 release). We also discuss the use of this generator in establishing a static benchmark suite to verify the correctness of programs submitted to the 2025 Google Code Golf Championship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。