对比两种方法在未知任务上的泛化能力,发现执行引导的程序合成更优。
Out-of-Distribution Generalization in the ARC-AGI Domain: Comparing Execution-Guided Neural Program Synthesis and Test-Time Fine-Tuning
- 用执行引导方式生成程序,提升组合新解题方案能力
- 在ARC-AGI任务中,新方法显著优于基准算法
- 适合关注模型泛化与推理机制的研究者
我们在ARC-AGI领域开展一项受控的组合泛化实验:这是一个开放世界问题域,泛化能力是成功的关键设计特征。我们比较了神经程序合成与测试时微调(TTFT)方法在此实验中的表现。结果表明,执行引导的神经程序合成在生成新颖解决方案方面优于所有参考算法。实证发现也显示,TTFT在ARC-AGI上的成功主要在于激发大语言模型原本无法直接依赖的分布内知识。
原文摘要 · Abstract (English)
We run a controlled compositional generalization experiment in the ARC-AGI domain: an open-world problem domain in which the ability to generalize out-of-distribution is, by design, an essential characteristic for success. We compare neural program synthesis and test-time fine-tuning approaches on this experiment. We find that execution-guided neural program synthesis outperforms all reference algorithms in its ability to compose novel solutions. Our empirical findings also suggest that the success of TTFT on ARC-AGI lies mainly in eliciting in-distribution knowledge that the LLM otherwise fails to rely on directly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。