用因果世界模型构建低数据推理测试集,验证大模型抽象推理能力
CausalARC: Abstract Reasoning with Causal World Models
- 基于结构因果模型生成可解释的推理任务
- 在少样本下实现观测、干预、反事实三种反馈学习
- 适合评估大模型在复杂逻辑与因果推理中的表现
实时推理常需在数据有限和分布漂移情况下适应新问题。本文提出CausalARC:一个模拟抽象与推理语料库(ARC)的实验测试平台,用于低数据和分布外场景下的人工智能推理研究。每个任务均来自完全定义的因果世界模型,以结构因果模型形式表达。通过合理数据增强,提供观测、干预和反事实反馈,以少样本、上下文学习的形式呈现。作为概念验证,我们展示了CausalARC在四种语言模型评估场景中的应用:(1) 测试时训练的抽象推理,(2) 基于上下文学习的反事实推理,(3) 程序合成,(4) 结合逻辑推理的因果发现。模型间与模型内性能在不同任务间差异显著,表明语言模型推理能力仍有巨大提升空间。
原文摘要 · Abstract (English)
On-the-fly reasoning often requires adaptation to novel problems under limited data and distribution shift. This work introduces CausalARC: an experimental testbed for AI reasoning in low-data and out-of-distribution regimes, modeled after the Abstraction and Reasoning Corpus (ARC). Each CausalARC reasoning task is sampled from a fully specified causal world model, formally expressed as a structural causal model. Principled data augmentations provide observational, interventional, and counterfactual feedback about the world model in the form of few-shot, in-context learning demonstrations. As a proof-of-concept, we illustrate the use of CausalARC for four language model evaluation settings: (1) abstract reasoning with test-time training, (2) counterfactual reasoning with in-context learning, (3) program synthesis, and (4) causal discovery with logical reasoning. Within- and between-model performance varied heavily across tasks, indicating room for significant improvement in language model reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。