LLM假设搜索比直接生成代码更接近人类推理能力
Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction
- 用假设搜索替代直接生成代码,提升少样本规则归纳表现
- 模型在测试中达到与人类相当的准确率,但直接生成差距明显
- 揭示了假设生成中的关键错误来源,为改进提供方向
归纳推理使人类能够从少量例子中推断出抽象规则并应用于新情境。本文比较了基于大语言模型(LLM)的假设搜索框架与直接程序生成方法在少样本规则归纳任务上的表现。结果表明,假设搜索的性能可媲美人类水平,而直接程序生成则显著落后。通过错误分析,我们识别出假设生成中的主要瓶颈,并提出了改进程序归纳方法的方向。总体而言,本工作凸显了基于LLM的假设搜索在建模归纳推理方面的潜力,以及构建更高效系统所面临的挑战。
原文摘要 · Abstract (English)
Inductive reasoning enables humans to infer abstract rules from limited examples and apply them to novel situations. In this work, we compare an LLM-based hypothesis search framework with direct program generation approaches on few-shot rule induction tasks. Our findings show that hypothesis search achieves performance comparable to humans, while direct program generation falls notably behind. An error analysis reveals key bottlenecks in hypothesis generation and suggests directions for advancing program induction methods. Overall, this paper underscores the potential of LLM-based hypothesis search for modeling inductive reasoning and the challenges in building more efficient systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。