arXiv:2510.20635cs.CLcs.AI2025-10ACL

用人类好奇心量表评估大模型,发现其求知欲更强但遇不确定仍保守。

Why Did Apple Fall: Evaluating Curiosity in Large Language Models

  • 基于五维好奇心量表设计评估框架,覆盖信息、刺激与社交三类好奇。
  • 大模型求知欲超人类,但在不确定情境下仍倾向保守决策。
  • 好奇行为能提升模型推理与主动学习能力,适合关注智能发展的研究者。

好奇心是人类探索与学习新知识的关键驱动力。近年来大语言模型(LLMs)在自然语言处理中的进展,引发了关于其是否具备类似人类的自主好奇学习能力的讨论。本文基于人类好奇心评估量表《五维好奇心量表修订版》(5DCR),构建了一个涵盖信息寻求、刺激寻求与社会好奇等维度的综合评估框架,用于量化评估大模型的好奇心表现。结果表明,大模型表现出比人类更强的知识渴求,但在面对不确定性环境时仍倾向于做出保守选择。进一步研究揭示,好奇行为可显著增强模型的推理能力与主动学习性能。这些发现为大模型未来实现类人好奇学习提供了实验支持,对推动其自主学习与创新研究具有重要意义。

原文摘要 · Abstract (English)

Curiosity serves as a pivotal conduit for human beings to discover and learn new knowledge. Recent advancements of large language models (LLMs) in natural language processing have sparked discussions regarding whether these models possess capability of curiosity-driven learning akin to humans. In this paper, starting from the human curiosity assessment questionnaire Five-Dimensional Curiosity scale Revised (5DCR), we design a comprehensive evaluation framework that covers dimensions such as Information Seeking, Thrill Seeking, and Social Curiosity to assess the extent of curiosity exhibited by LLMs. The results demonstrate that LLMs exhibit a stronger thirst for knowledge than humans but still tend to make conservative choices when faced with uncertain environments. We further investigated the relationship between curiosity and thinking of LLMs, confirming that curious behaviors can enhance the model's reasoning and active learning abilities. These findings suggest that LLMs have the potential to exhibit curiosity similar to that of humans, providing experimental support for the future development of learning capabilities and innovative research in LLMs.

大模型好奇心评估推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。