arXiv:2502.07443cs.AIcs.GT2025-02被引 8

用LLM增强的多智能体框架模拟人类策略推理,表现优于传统模型。

Approximating Human Strategic Reasoning with LLM-Enhanced Recursive Reasoners Leveraging Multi-agent Hypergames

  • 基于超博弈的层级信念模型,实现角色化多智能体策略互动
  • 在单次两人猜谜游戏中,人工推理者逼近人类行为并达最优解
  • 适合研究博弈论、社会模拟与大模型战略能力的学者

基于LLM的多智能体仿真在博弈论与社会模拟中日益流行。现有方法多采用弱代理概念和简化架构。本文构建了面向复杂递归推理的角色化多智能体框架,支持系统性开发与评估策略推理。游戏由裁判负责匹配、走步验证与环境管理。玩家采用前沿大语言模型,依赖形式化的超博弈层级信念模型进行决策。我们通过单次、两人制猜谜游戏评估最新LLMs的递归推理能力,并与经济学基准模型及人类实验数据对比。此外,提出一种新的语义化推理度量方式,用于k级理论分析。实验表明,人工推理者在逼近人类行为与达成最优解方面均优于基准模型。

原文摘要 · Abstract (English)

LLM-driven multi-agent-based simulations have been gaining traction with applications in game-theoretic and social simulations. While most implementations seek to exploit or evaluate LLM-agentic reasoning, they often do so with a weak notion of agency and simplified architectures. We implement a role-based multi-agent strategic interaction framework tailored to sophisticated recursive reasoners, providing the means for systematic in-depth development and evaluation of strategic reasoning. Our game environment is governed by the umpire responsible for facilitating games, from matchmaking through move validation to environment management. Players incorporate state-of-the-art LLMs in their decision mechanism, relying on a formal hypergame-based model of hierarchical beliefs. We use one-shot, 2-player beauty contests to evaluate the recursive reasoning capabilities of the latest LLMs, providing a comparison to an established baseline model from economics and data from human experiments. Furthermore, we introduce the foundations of an alternative semantic measure of reasoning to the k-level theory. Our experiments show that artificial reasoners can outperform the baseline model in terms of both approximating human behaviour and reaching the optimal solution.

策略推理多智能体超博弈LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。