arXiv:2510.10182cs.CLcs.AI2025-10ACL综述被引 7

首篇系统综述大模型归纳推理,梳理方法与评测体系。

A Survey of Inductive Reasoning for Large Language Models

  • 按后训练、测试时扩展、数据增强三类整理提升归纳推理的方法
  • 构建统一沙盒评估框架,引入观察覆盖率指标量化表现
  • 揭示简单架构与数据对归纳能力的促进作用,适合研究者参考

推理是大语言模型的重要任务。在各类推理范式中,归纳推理是一种基础类型,其特征为从特定案例推导一般规律,且答案不唯一。该模式有利于知识泛化,更符合人类认知,是学习的核心方式,因而日益受到关注。尽管重要,归纳推理尚无系统总结。本文首次全面综述大模型的归纳推理,将提升方法分为后训练、测试时扩展和数据增强三类;总结现有基准并提出基于沙盒的统一评估方法,引入观察覆盖率指标;最后分析归纳能力来源,揭示简单模型架构与数据对归纳任务的积极作用,为未来研究奠定基础。

原文摘要 · Abstract (English)

Reasoning is an important task for large language models (LLMs). Among all the reasoning paradigms, inductive reasoning is one of the fundamental types, which is characterized by its particular-to-general thinking process and the non-uniqueness of its answers. The inductive mode is crucial for knowledge generalization and aligns better with human cognition, so it is a fundamental mode of learning, hence attracting increasing interest. Despite the importance of inductive reasoning, there is no systematic summary of it. Therefore, this paper presents the first comprehensive survey of inductive reasoning for LLMs. First, methods for improving inductive reasoning are categorized into three main areas: post-training, test-time scaling, and data augmentation. Then, current benchmarks of inductive reasoning are summarized, and a unified sandbox-based evaluation approach with the observation coverage metric is derived. Finally, we offer some analyses regarding the source of inductive ability and how simple model architectures and data help with inductive tasks, providing a solid foundation for future research.

归纳推理大模型综述评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。