arXiv:2605.17278cs.AIcs.LG2026-05被引 1

用自动验证方法生成可信赖的抽象推理测试集

A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation

论文配图:A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
图 1 · 摘自论文原文
  • 用大模型自动生成需抽象推理的任务并扩展
  • 实测顶级模型在3D任务上仅达人类39.8%
  • 发现复杂输入反而简化推理过程

抽象推理能力反映大模型提取并应用抽象规则的智能水平。然而,准确衡量该能力仍具挑战:现有基准或依赖昂贵的人工标注,限制规模;或可能测量记忆而非真实推理。为此,我们提出自动化流程A2RBench,涵盖生成、扩展、评估与分析。生成阶段,大模型创建多样化需真实推理的任务;扩展阶段,复用已验证规则并拓展新输入空间,实现规模扩展。但此过程可能引发幻觉。为此,我们建立理论框架,证明程序化验证——检验逆操作能否完美逆转正向操作(循环一致性)——可保证唯一解。在主流大模型上的广泛评估显示:(1) 当前大模型在抽象推理上存在根本缺陷,顶尖模型在代表性子集上表现远低于人类(39.8% vs. 68.5%);(2) 在生成的3D任务复杂度上,当前大模型显著落后于2D和1D任务,暴露其对高维任务理解不足;(3) 反直觉地,信息复杂度更高的输入反而能简化推理过程。

原文摘要 · Abstract (English)

Abstract reasoning ability reflects the intelligence and generalization capacity of LLMs to extract and apply abstract rules. However, accurately measuring this ability remains challenging: existing benchmarks either rely on expensive manual annotation, limiting their scale, or risk measuring memorization rather than genuine reasoning. To address this, we introduce an automated pipeline named A2RBench, encompassing generation, expansion, evaluation, and analysis. Specifically, in the generation stage, LLMs create diverse tasks demanding genuine reasoning; in the expansion stage, LLMs reuse validated rules and expand new input spaces to generate task variations, achieving scaling. However, such a process may cause hallucinations. To eliminate it, we further establish a theoretical framework and prove that programmatic verification--testing whether the inverse operation perfectly reverses the forward operation (cycle consistency)--guarantees a unique solution. Through extensive evaluations on mainstream LLMs, we find: (1) Current LLMs exhibit fundamental deficiencies in abstract reasoning, with top models significantly underperforming humans on a representative subset (39.8% vs. 68.5%). (2) Current LLMs fall far short of 2D and 1D in the complexity of generated 3D tasks, revealing their lack of understanding of high-dimensional tasks. (3) Counterintuitively, inputs with higher information complexity can simplify the reasoning process.

抽象推理自动评测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。