让大模型在无公式场景中像科学家一样归纳规律,挑战推理新高度。
On LLM-Based Scientific Inductive Reasoning Beyond Equations
- 基于科学发现类比,设计非方程类归纳任务
- 构建首个面向科学场景的归纳推理基准SIRBench-V1
- 揭示当前大模型在新环境下的推理短板
随着大语言模型(LLMs)日益展现类人能力,一个核心问题浮现:如何使它们在全新环境中仅凭少量示例学习潜在规律并有效应用?这正是归纳推理的关键。现有研究多按规则是否可由显式数学方程表达进行分类,但许多‘超越方程’类研究缺乏具体场景支撑。受人类科学发现过程启发,我们提出‘大模型科学归纳推理’任务,并引入新基准SIRBench-V1,用于评估模型在科学情境下的归纳能力。实验表明,当前大模型在此任务上仍表现不佳,凸显其挑战性与进一步发展的必要性。
原文摘要 · Abstract (English)
As large language models (LLMs) increasingly exhibit human-like capabilities, a fundamental question emerges: How can we enable LLMs to learn the underlying patterns from limited examples in entirely novel environments and apply them effectively? This question is central to the ability of LLMs in inductive reasoning. Existing research on LLM-based inductive reasoning can be broadly categorized based on whether the underlying rules are expressible via explicit mathematical equations. However, many recent studies in the beyond-equations category have emphasized rule design without grounding them in specific scenarios. Inspired by the parallels between inductive reasoning and human scientific discovery, we propose the task of LLM-Based Scientific Inductive Reasoning Beyond Equations and introduce a new benchmark, SIRBench-V1, to evaluate the inductive reasoning abilities of LLMs in scientific settings. Our experimental results show that current LLMs still struggle with this task, underscoring its difficulty and the need for further advancement in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。