arXiv:2608.16196cs.AIcs.HC2026-08

用行为数据精准推断玩家能力与风格,让游戏自动个性化适配。

Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior

论文配图:Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior
图 1 · 摘自论文原文
  • 构建合成玩家群体,以可验证的参数作为真实行为基准。
  • 提出时机感知的行为分析方法,区分偏好与表达机会。
  • 验证大模型可有效推断玩家特质,支持动态难度调节。

个性化游戏生成需从玩家行为中推断其能力与风格。大型语言模型可将原始游戏日志转化为流畅且合理的玩家画像,但这些画像缺乏验证。本文提出一个合成玩家群体,其属性由明确的机器人参数定义,确保行为变化与特定属性一致。该基准不依赖已知决策模型,仅基于行为日志进行无模型推断。引入机会感知的决策时刻表征,分离偏好与表达机会,消融实验显示其对依赖机会的属性影响显著。在该基准上,少样本大模型推理优于嵌入与规则基线,但特征驱动的监督回归仍更优。最后,将推断结果用于动态难度调节,对照真实标签与错误配置,结合初步人类实验验证迁移可行性。

原文摘要 · Abstract (English)

Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plausible profile of the player. Plausible, however, is not verified, and verification is precisely what the field lacks: latent traits are unobservable; questionnaires provide noisy proxies and become circular when self-reports are used to validate behavior-based inference; and behavior itself is ambiguous without context -- a player who never collects an item may not want it, or may never have had the chance. We address both problems. First, we construct a synthetic player population whose traits are ground truth by construction: each trait is an explicit bot parameter, accepted only after controlled manipulation produces consistent, trait-specific behavioral change. Unlike prior parameter-recovery work that inverts a known decision model, our benchmark evaluates policy-agnostic inference from behavioral transcripts alone. Second, we introduce an opportunity-aware decision-moment representation that disentangles preference from the chance to express it; ablating it selectively degrades opportunity-dependent traits. On this benchmark, few-shot LLM inference outperforms embedding- and rule-based baselines on most traits, though feature-based supervised regressors remain stronger overall. Finally, we close the loop: inferred profiles drive difficulty adaptation, evaluated against ground-truth references and mismatched-profile controls, and an exploratory human study examines whether these findings transfer to real players.

游戏生成行为推断大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。