用小模型实现空间组合推理的系统泛化,效果超大模型。
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
- 通过元学习训练小模型掌握几何变换组合能力。
- 570万参数模型在未见组合上表现超越GPT-4o等大模型。
- 为非语言任务提升系统泛化提供新思路,适合模型泛化研究者。
系统泛化指从已知组件中理解并生成新组合的能力。尽管大语言模型在多领域取得进展,但在面对新组合场景时常表现不佳,暴露出系统泛化能力的不足。学界对神经网络是否具备系统泛化能力存在争议,近期研究表明针对组合性的元学习方法可显著提升该能力,但相关成果主要局限于语言任务,其在其他任务中的适用性尚不明确。本文将组合性元学习拓展至抽象空间推理领域,提出Compositional-ARC数据集,用于评估模型从已知几何变换(如平移、旋转)到新组合(如平移+旋转)的系统泛化能力。实验表明,一个仅含570万参数的Transformer编码器-解码器模型,经组合性元学习训练后,能有效泛化至未见过的变换组合。值得注意的是,该模型性能显著优于o3-mini、GPT-4o和Gemini 2.0 Flash等前沿大模型,且与2024年ARC竞赛冠军(80亿参数大模型,通过测试时训练获得)表现相当。结果表明,元学习在非语言任务中亦能有效促进系统性,为构建更鲁棒、通用的模型提供了新方向。
原文摘要 · Abstract (English)
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language models (LLMs) across various domains, these models often fail to extend their knowledge to novel compositional scenarios, revealing notable limitations in systematic generalization. There has been an ongoing debate about whether neural networks possess the capacity for systematic generalization, with recent studies suggesting that meta-learning approaches designed for compositionality can significantly enhance this ability. However, these insights have largely been confined to linguistic problems, leaving their applicability to other tasks an open question. In this study, we extend meta-learning for compositionality to the domain of abstract spatial reasoning. To this end, we introduce $\textit{Compositional-ARC}\unicode{x2014}$a dataset designed to evaluate the capacity of models to systematically generalize from known geometric transformations (e.g., translation, rotation) of abstract two-dimensional objects to novel combinations of these transformations (e.g., translation+rotation). Our results show that a small transformer-based encoder-decoder model, trained via meta-learning for compositionality, can systematically generalize to previously unseen transformation compositions. Notably, despite having only 5.7M parameters, this model significantly outperforms state-of-the-art LLMs$\unicode{x2014}$including o3-mini, GPT-4o, and Gemini 2.0 Flash, which fail to exhibit similar systematic behavior$\unicode{x2014}$and performs on par with the winning model of the ARC prize 2024, an 8B-parameter LLM trained via test-time training. Our findings highlight the effectiveness of meta-learning in promoting systematicity beyond linguistic tasks, suggesting a promising direction toward more robust and generalizable models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。