arXiv:2510.06151cs.LGcs.AI2025-10被引 1

用大模型模拟人类队友,低成本构建可扩展的异构智能体协作环境。

LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams

  • 用大模型作为无策略依赖的人类代理,通过提示词生成模仿人类决策的数据。
  • 大模型在博弈任务中决策与专家一致,且能根据提示调整风险偏好,匹配人类行为多样性。
  • 适合需要模拟真实人类协作的强化学习研究,尤其在缺乏真人数据时使用。

在异构智能体团队中,训练智能体与策略不可见或非稳态的队友(如人类)协作是一项关键挑战。传统方法依赖昂贵的人类参与数据,限制了可扩展性。本文提出使用大语言模型(LLMs)作为无策略依赖的人类代理,生成模拟人类决策的合成数据。我们在受狩猎游戏启发的网格世界任务中开展三项实验:实验一将30名人类参与者和2位专家的决策与LLaMA 3.1及Mixtral 8x22B模型输出对比,发现模型在给定游戏状态和奖励结构下,比人类更接近专家判断,表现出一致的决策逻辑;实验二通过提示词引导,使模型展现出风险规避或风险追求等多样化策略,与人类参与者的行为变异一致;实验三在动态网格世界中,模型生成的动作轨迹与人类路径高度相似。尽管大模型尚无法完全复现人类的适应性,但其提示驱动的多样性为模拟无策略依赖队友提供了可扩展的基础。

原文摘要 · Abstract (English)

A critical challenge in modelling Heterogeneous-Agent Teams is training agents to collaborate with teammates whose policies are inaccessible or non-stationary, such as humans. Traditional approaches rely on expensive human-in-the-loop data, which limits scalability. We propose using Large Language Models (LLMs) as policy-agnostic human proxies to generate synthetic data that mimics human decision-making. To evaluate this, we conduct three experiments in a grid-world capture game inspired by Stag Hunt, a game theory paradigm that balances risk and reward. In Experiment 1, we compare decisions from 30 human participants and 2 expert judges with outputs from LLaMA 3.1 and Mixtral 8x22B models. LLMs, prompted with game-state observations and reward structures, align more closely with experts than participants, demonstrating consistency in applying underlying decision criteria. Experiment 2 modifies prompts to induce risk-sensitive strategies (e.g. "be risk averse"). LLM outputs mirror human participants' variability, shifting between risk-averse and risk-seeking behaviours. Finally, Experiment 3 tests LLMs in a dynamic grid-world where the LLM agents generate movement actions. LLMs produce trajectories resembling human participants' paths. While LLMs cannot yet fully replicate human adaptability, their prompt-guided diversity offers a scalable foundation for simulating policy-agnostic teammates.

大模型协作智能体人类模拟强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。