arXiv:2504.11671cs.AIcs.CY2025-04被引 2

通过操控模型内部向量,揭示并调节大模型在社会博弈中的决策机制。

Computational Basis of LLM's Decision Making in Social Simulation

  • 提取模型内部的变量变化向量,实现对社会概念编码的可操控性。
  • 改变性别等变量向量能显著影响模型在公平博弈中的分配行为。
  • 为社会模拟中的AI代理对齐与去偏提供可解释的干预方法。

大语言模型(LLMs)在社会科学与实际应用中越来越多地作为类人决策代理使用。这些模型代理通常被赋予类人角色并置于真实情境中,但角色与情境如何塑造其行为仍缺乏深入研究。本研究提出并验证了在经典公平博弈实验——独裁者游戏(Dictator Game)中,探测、量化和修改LLM内部表征的方法。通过从模型内部状态中提取“变量变化向量”(如从‘男性’到‘女性’),在推理过程中操纵这些向量,可显著改变变量与模型决策之间的关联。该方法为研究和调控社会概念在基于Transformer的模型中的编码与工程实现提供了系统性路径,对对齐、去偏以及学术与商业场景中社会模拟代理的设计具有重要意义,有助于加强社会学理论与测量。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly serve as human-like decision-making agents in social science and applied settings. These LLM-agents are typically assigned human-like characters and placed in real-life contexts. However, how these characters and contexts shape an LLM's behavior remains underexplored. This study proposes and tests methods for probing, quantifying, and modifying an LLM's internal representations in a Dictator Game, a classic behavioral experiment on fairness and prosocial behavior. We extract ``vectors of variable variations'' (e.g., ``male'' to ``female'') from the LLM's internal state. Manipulating these vectors during the model's inference can substantially alter how those variables relate to the model's decision-making. This approach offers a principled way to study and regulate how social concepts can be encoded and engineered within transformer-based models, with implications for alignment, debiasing, and designing AI agents for social simulations in both academic and commercial applications, strengthening sociological theory and measurement.

大模型决策社会模拟可解释性去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。