用猜谜游戏测试大模型智能体,发现模仿人类思维能提升表现但不简单叠加。
The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners
- 设计三类智能体:传统博弈模型、直接调用大模型、结合传统框架的大模型
- 2000+样本显示,模拟人类认知结构的智能体更接近人类策略行为
- 智能体复杂度与人类相似性非线性相关,依赖底层大模型能力
大型语言模型(LLM)的快速发展推动了智能体系统的发展,促使研究采用更弱且更灵活的代理概念。然而,这引发了关键问题:基于LLM的智能体在博弈论场景中能否复现人类的战略推理?本文通过猜谜游戏作为测试平台,评估三种智能体设计:简单的博弈论模型、无结构的LLM作为代理、以及将LLM整合到传统代理框架中的模型。这些智能体与人类参与者在一般推理模式和个体角色目标上进行对比。此外,引入模糊化游戏场景以评估智能体在训练分布之外的泛化能力。分析覆盖25种智能体配置下的2000多个推理样本,结果表明,受人类启发的认知结构可增强LLM智能体与人类战略行为的一致性。然而,智能体设计复杂度与人类相似性之间呈现非线性关系,凸显对底层LLM能力的高度依赖,并暗示单纯架构改进存在局限。
原文摘要 · Abstract (English)
The rapid rise of large language models (LLMs) has shifted artificial intelligence (AI) research toward agentic systems, motivating the use of weaker and more flexible notions of agency. However, this shift raises key questions about the extent to which LLM-based agents replicate human strategic reasoning, particularly in game-theoretic settings. In this context, we examine the role of agentic sophistication in shaping artificial reasoners' performance by evaluating three agent designs: a simple game-theoretic model, an unstructured LLM-as-agent model, and an LLM integrated into a traditional agentic framework. Using guessing games as a testbed, we benchmarked these agents against human participants across general reasoning patterns and individual role-based objectives. Furthermore, we introduced obfuscated game scenarios to assess agents' ability to generalise beyond training distributions. Our analysis, covering over 2000 reasoning samples across 25 agent configurations, shows that human-inspired cognitive structures can enhance LLM agents' alignment with human strategic behaviour. Still, the relationship between agentic design complexity and human-likeness is non-linear, highlighting a critical dependence on underlying LLM capabilities and suggesting limits to simple architectural augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。