arXiv:2603.21398cs.AIcs.GT2026-03被引 5

用激活向量测量并操控大模型在博弈中的性格特征。

Persona Vectors in Games: Measuring and Steering Strategies via Activation Vectors

  • 通过对比激活叠加构建利他、宽容等性格向量。
  • 操纵向量后策略与语言解释出现系统性变化。
  • 适合研究模型行为机制或可控生成的学者。

大型语言模型(LLMs)越来越多地被用作战略环境中的自主决策者,但我们缺乏理解其高层次行为特征的工具。本文在博弈论场景中使用激活转向方法,通过对比激活叠加构建利他、宽容及对他人的期望等人格向量。在经典博弈任务上评估发现,激活转向能系统性改变策略选择和自然语言解释。然而也观察到,在转向下,言辞与策略可能出现分离。此外,自我行为与对他人预期的向量部分独立。结果表明,人格向量为战略环境中高层特质提供了有前景的机制化控制手段。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as autonomous decision-makers in strategic settings, yet we have limited tools for understanding their high-level behavioral traits. We use activation steering methods in game-theoretic settings, constructing persona vectors for altruism, forgiveness, and expectations of others by contrastive activation addition. Evaluating on canonical games, we find that activation steering systematically shifts both quantitative strategic choices and natural-language justifications. However, we also observe that rhetoric and strategy can diverge under steering. In addition, vectors for self-behavior and expectations of others are partially distinct. Our results suggest that persona vectors offer a promising mechanistic handle on high-level traits in strategic environments.

人格向量博弈论激活控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。