arXiv:2512.20162cs.AI2025-12被引 1

对比人与大模型在数列推理中的概念泛化能力,发现人类更灵活,大模型依赖数学规则。

Concept Generalization in Humans and Large Language Models: Insights from the Number Game

  • 用贝叶斯模型分析人类与大模型的归纳偏置和推理策略。
  • 人类仅需一个例子即可泛化,大模型需更多样本才能完成。
  • 揭示人类兼具规则与相似性推理,大模型更依赖数学规律。

我们比较了人类与大型语言模型(LLM)在数列游戏这一概念推断任务中的泛化能力。基于贝叶斯模型的分析框架,研究了人类与大模型的归纳偏置和推理策略。结果表明,贝叶斯模型对人类行为的拟合优于对大模型的表现:人类能灵活地推断基于规则和基于相似性的概念,而大模型则更依赖数学规则。此外,人类表现出少样本泛化能力,甚至可从单一示例中学习,而大模型需要更多样本才能实现泛化。这些差异凸显了人类与大模型在数学概念推断与泛化机制上的根本不同。

原文摘要 · Abstract (English)

We compare human and large language model (LLM) generalization in the number game, a concept inference task. Using a Bayesian model as an analytical framework, we examined the inductive biases and inference strategies of humans and LLMs. The Bayesian model captured human behavior better than LLMs in that humans flexibly infer rule-based and similarity-based concepts, whereas LLMs rely more on mathematical rules. Humans also demonstrated a few-shot generalization, even from a single example, while LLMs required more samples to generalize. These contrasts highlight the fundamental differences in how humans and LLMs infer and generalize mathematical concepts.

概念泛化人类认知大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。