arXiv:2505.17482cs.AIcs.CL2025-05被引 2

用知识增强提升大模型在抽象推理任务中的泛化能力

From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark

  • 构建分层知识体系,逐步注入先验知识以增强推理
  • 在ARC基准上实现平均5%绝对提升,最高达64.52%相对增益
  • 适合关注模型泛化与认知推理能力的研究者

近期面向推理的大语言模型在数学和科学考试等挑战性任务中表现优异,但人类智能的核心能力如抽象推理与泛化仍缺乏深入探索。为此,我们评估了现有推理型大模型在抽象与推理语料库(ARC)基准上的表现,该基准明确要求具备这两种能力。我们将ARC建模为程序合成任务,并提出九种候选求解器。实验表明,重复采样规划辅助代码生成(RSPC)在测试准确率上表现最佳,并在多数大模型中展现出一致的泛化能力。为进一步提升性能,我们提出一种名为知识增强抽象推理(KAAR)的求解器,其通过一个分类为三个层级的本体结构编码核心知识先验。KAAR逐级扩展模型推理能力,每阶段后调用RSPC生成候选解。这种分阶段推理有效减少无关先验干扰,显著提升模型表现。实证结果表明,KAAR在所有评估的大模型上均持续优于非增强版RSPC,实现约5%的绝对提升,最高达64.52%的相对改进。尽管取得进展,ARC对推理型大模型而言仍是极具挑战的基准,凸显了未来大模型发展的关键方向。

原文摘要 · Abstract (English)

Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus (ARC) benchmark, which explicitly demands both faculties. We formulate ARC as a program synthesis task and propose nine candidate solvers. Experimental results show that repeated-sampling planning-aided code generation (RSPC) achieves the highest test accuracy and demonstrates consistent generalization across most LLMs. To further improve performance, we introduce an ARC solver, Knowledge Augmentation for Abstract Reasoning (KAAR), which encodes core knowledge priors within an ontology that classifies priors into three hierarchical levels based on their dependencies. KAAR progressively expands LLM reasoning capacity by gradually augmenting priors at each level, and invokes RSPC to generate candidate solutions after each augmentation stage. This stage-wise reasoning reduces interference from irrelevant priors and improves LLM performance. Empirical results show that KAAR maintains strong generalization and consistently outperforms non-augmented RSPC across all evaluated LLMs, achieving around 5% absolute gains and up to 64.52% relative improvement. Despite these achievements, ARC remains a challenging benchmark for reasoning-oriented LLMs, highlighting future avenues of progress in LLMs.

抽象推理知识增强大模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。