arXiv:2509.05249cs.CVcs.AI2025-09

构建视觉推理框架COGITAO,测试模型组合与泛化能力

COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization

  • 基于规则生成任务,支持28种可组合变换
  • 生成数百万独特任务,难度跨度大且可无限采样
  • 揭示现有视觉模型在新组合下严重泛化失败

人类智能的核心在于能组合已学概念并在新情境中应用,而当前顶尖机器学习模型仍难以实现这一能力。为此,我们提出COGITAO——一个模块化、可扩展的数据生成框架与基准,系统研究视觉领域的组合性与泛化能力。受ARC-AGI问题设定启发,COGITAO在网格环境中通过一系列规则化变换操作物体,支持28种可互操作的变换,并可在可调深度上实现组合;同时对网格参数与物体属性有精细控制。该设计使框架可生成数百万种独特任务规则,远超现有数据集数量级,覆盖广泛难度,且每条规则可无限生成样本。我们使用前沿视觉模型进行基线实验,发现其虽在域内表现良好,却持续无法泛化到熟悉元素的新组合。COGITAO已全面开源,包含全部代码与数据集,以推动该方向持续研究。

原文摘要 · Abstract (English)

The ability to compose learned concepts and apply them in novel settings is key to human intelligence, but remains a persistent limitation in state-of-the-art machine learning models. To address this issue, we introduce COGITAO, a modular and extensible data generation framework and benchmark designed to systematically study compositionality and generalization in visual domains. Drawing inspiration from ARC-AGI's problem-setting, COGITAO constructs rule-based tasks which apply a set of transformations to objects in grid-like environments. It supports composition, at adjustable depth, over a set of 28 interoperable transformations, along with extensive control over grid parametrization and object properties. This flexibility enables the creation of millions of unique task rules -- surpassing concurrent datasets by several orders of magnitude -- across a wide range of difficulties, while allowing virtually unlimited sample generation per rule. We provide baseline experiments using state-of-the-art vision models, highlighting their consistent failures to generalize to novel combinations of familiar elements, despite strong in-domain performance. COGITAO is fully open-sourced, including all code and datasets, to support continued research in this field.

视觉推理组合性泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。