arXiv:2603.03295cs.CLcs.AI2026-03

LLM选目标能力远不如人,自导向学习中表现差异大。

Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

  • 用认知科学实验任务测试模型自主选目标能力
  • 多数模型只依赖单一解法,表现远低于人类
  • 不同模型差异显著,适合研究人类决策机制

大型语言模型(LLMs)在代理工作流、社交研究或聊天场景中被越来越多地要求自主选择目标,而非完成预设任务。然而,当前假设认为LLMs能准确反映人类目标偏好,尚缺乏验证。本文在借鉴认知科学的自导向学习任务中,评估了五种模型(GPT-5、Gemini 2.5 Pro、Claude Sonnet 4.5、Qwen3 32B、Centaur)作为人类目标选择代理的有效性。结果发现,人类会逐步探索并多样化达成目标,而多数模型仅依赖单一解法,表现明显偏低,且不同模型间模式各异,同模型内部变异性低。链式思维与人格引导仅带来有限改进,结论在不同实验设置下均成立。这些发现提示当前模型无法替代人类的目标选择行为,强调其独特性,需谨慎用于实际应用。

原文摘要 · Abstract (English)

Whether in agentic workflows, social studies, or chat settings, large language models (LLMs) are increasingly being asked to replace humans in choosing which goals to pursue, rather than completing predefined tasks. However, the assumption that LLMs accurately reflect human preferences for goal setting remains largely untested. We assess the validity of LLMs as proxies for human goal selection in a controlled, self-directed learning task borrowed from cognitive science. Across five models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5, Qwen3 32B, and Centaur), we find substantial divergence from human behavior. While people gradually explore and learn to achieve goals with diversity across individuals, most models exploit a single identified solution or show surprisingly low performance, with distinct patterns across models and little variability across instances of the same model. Chain-of-thought reasoning and persona steering provide limited improvements, and our conclusions hold across experimental settings. While they await confirmation in applied settings, these findings highlight the uniqueness of human goal selection and caution against its replacement with current models.

目标选择语言模型认知科学自导向学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。