检验大模型角色扮演时说的和做的是否一致,发现常有偏差。
Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
- 通过信任游戏测试大模型角色信念与行为的一致性
- 即使信念看似合理,行为仍可能不一致,尤其在远期预测时
- 提醒研究者注意大模型生成数据的可信度,适合行为模拟研究
随着大模型被用于角色扮演以生成人类行为研究的合成数据,确保其输出与所分配角色一致成为关键问题。本文探究大模型在角色扮演中所述信念(‘怎么说’)与其实际行为(‘怎么做’)的一致性。我们建立评估框架,衡量通过提示获取的信念能否提前预测模拟结果。基于增强版GenAgents人物库和信任游戏(一种标准经济博弈,用于量化信任与互惠),引入信念-行为一致性指标,系统考察了三类因素的影响:(1)获取信念的方式,如模拟预期结果或角色任务相关属性;(2)向模型提供信任游戏信息的时间与方式;(3)要求模型预测未来行为的时间跨度。我们还探索在原始信念与研究目标不符时,强行施加理论先验的可行性。结果表明,大模型的陈述信念与模拟行为之间存在系统性不一致,无论个体还是群体层面均如此。即使模型表现出合理信念,也可能无法持续应用。这凸显了识别信念与行为对齐时机的重要性,以便研究人员恰当使用大模型进行行为研究。
原文摘要 · Abstract (English)
As LLMs are increasingly studied as role-playing agents to generate synthetic data for human behavioral research, ensuring that their outputs remain coherent with their assigned roles has become a critical concern. In this paper, we investigate how consistently LLM-based role-playing agents' stated beliefs about the behavior of the people they are asked to role-play ("what they say") correspond to their actual behavior during role-play ("how they act"). Specifically, we establish an evaluation framework to rigorously measure how well beliefs obtained by prompting the model can predict simulation outcomes in advance. Using an augmented version of the GenAgents persona bank and the Trust Game (a standard economic game used to quantify players' trust and reciprocity), we introduce a belief-behavior consistency metric to systematically investigate how it is affected by factors such as: (1) the types of beliefs we elicit from LLMs, like expected outcomes of simulations versus task-relevant attributes of individual characters LLMs are asked to simulate; (2) when and how we present LLMs with relevant information about Trust Game; and (3) how far into the future we ask the model to forecast its actions. We also explore how feasible it is to impose a researcher's own theoretical priors in the event that the originally elicited beliefs are misaligned with research objectives. Our results reveal systematic inconsistencies between LLMs' stated (or imposed) beliefs and the outcomes of their role-playing simulation, at both an individual- and population-level. Specifically, we find that, even when models appear to encode plausible beliefs, they may fail to apply them in a consistent way. These findings highlight the need to identify how and when LLMs' stated beliefs align with their simulated behavior, allowing researchers to use LLM-based agents appropriately in behavioral studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。