arXiv:2608.07490cs.HCcs.AI2026-08

对比人类与语言智能体在反复对战中的行为演化,发现当前智能体进化不持久。

Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents

论文配图:Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents
图 1 · 摘自论文原文
  • 设计可复用策略结构的游戏框架,通过动作轨迹分析行为变化
  • 人类玩家从局部贪心转向全局策略,行为演变稳定可解释
  • 现有自进化智能体表现噪声大、效果短暂,难持续优化决策

大型语言模型智能体正越来越多地通过游戏进行评估,但多数基准仅关注最终结果而非学习过程。本文研究经验敏感型游戏学习:即重复对战如何改变人类与语言智能体的决策行为。我们提出一个分析行为随游戏经验演化的框架,超越单纯依赖最终得分或胜率。构建了一套具有可复用战略结构的互动游戏,并引入跨游戏的贪婪到全局度量及特定游戏的行为诊断工具,使从动作轨迹中观测经验驱动的变化成为可能。同时收集了人类玩家的重复对战轨迹,并在相同行为度量空间下评估近期自进化语言智能体。结果表明,人类玩家表现出可解释且相对稳定的转变,从局部贪心启发式逐步转向更全局的战略决策;而当前自进化智能体则常呈现噪声大、短暂的收益,说明现有自进化方法仍难以将游戏经验转化为持久的决策行为改变。

原文摘要 · Abstract (English)

Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from locally greedy heuristics toward more global strategic decisions. In contrast, current self-evolving agents often show noisy and transient gains, suggesting that existing self-evolution methods remain limited in converting gameplay experience into durable changes in decision-making behavior.

智能体学习行为演化游戏评估语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。