对比人类与大模型生成任务,发现大模型缺乏真实动机和身体体验。
Mind the Gap: The Divergence Between Human and LLM-Generated Tasks
- 用心理因素引导人类和大模型生成任务,对比行为差异。
- 大模型生成任务更抽象、少社交、少身体活动,但被认为更有趣新颖。
- 适合关注人机认知差异与具身智能的研究者参考。
人类在内在动机驱动下持续生成多样化任务。尽管基于大语言模型(LLMs)的生成代理旨在模拟这种复杂行为,但其是否遵循相似认知机制仍不明确。为此,我们开展了一项任务生成实验,比较人类与一个大模型代理(GPT-4o)的表现。结果表明,人类任务生成始终受心理驱动影响,如个人价值观(例如对变化的开放性)和认知风格。即使将这些心理因素显式提供给大模型,其仍未能体现相应的行为模式,生成的任务显著更少社会性、更少身体参与,且主题更偏向抽象。有趣的是,大模型生成的任务虽被感知为更具趣味性和新颖性,这凸显了其语言能力与生成具身化目标能力之间的脱节。结论表明,人类认知的核心特征——价值驱动与具身性——与大模型的统计模式之间存在根本差距,强调在构建更符合人性的代理时,需融入内在动机与物理基础。
原文摘要 · Abstract (English)
Humans constantly generate a diverse range of tasks guided by internal motivations. While generative agents powered by large language models (LLMs) aim to simulate this complex behavior, it remains uncertain whether they operate on similar cognitive principles. To address this, we conducted a task-generation experiment comparing human responses with those of an LLM agent (GPT-4o). We find that human task generation is consistently influenced by psychological drivers, including personal values (e.g., Openness to Change) and cognitive style. Even when these psychological drivers are explicitly provided to the LLM, it fails to reflect the corresponding behavioral patterns. They produce tasks that are markedly less social, less physical, and thematically biased toward abstraction. Interestingly, while the LLM's tasks were perceived as more fun and novel, this highlights a disconnect between its linguistic proficiency and its capacity to generate human-like, embodied goals. We conclude that there is a core gap between the value-driven, embodied nature of human cognition and the statistical patterns of LLMs, highlighting the necessity of incorporating intrinsic motivation and physical grounding into the design of more human-aligned agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。