通过构建情境与导航推理,提升大模型在复杂决策中的行为一致性。
Improving Behavioral Alignment in LLM Social Simulations via Context Formation and Navigation
- 分两阶段:先建精准决策情境,再引导模型在其中推理决策。
- 四款主流模型在复杂任务中需双阶段才能对齐人类行为,简单任务仅需第一阶段。
- 适用于需要高仿真人类行为的实验心理学与社会模拟研究。
大语言模型在模拟人类行为时,常在复杂决策环境中偏离真实人类决策,这类环境要求参与者预判他人行动并基于观察形成信念。本文提出两阶段框架:第一阶段为情境构建,明确实验设计以准确表征决策任务及其背景;第二阶段为情境导航,在该表征内引导模型推理以做出决策。通过复现一个带有质量信号的顺序购买游戏(Kremer and Debo, 2016),并扩展至包含成本信号的众筹游戏(Cason et al., 2025)和需求估计任务(Gui and Toubia, 2025),验证了该框架在多种决策环境下的泛化能力。在GPT-4o、GPT-5、Claude-4.0-Sonnet-Thinking、DeepSeek-R1四款SOTA模型上,复杂任务需双阶段才可实现与人类基准的行为对齐,而简单的需求估计任务仅需情境构建即可。研究明确了各阶段适用条件,为设计与诊断大模型社会模拟提供了系统方法,使其成为行为研究中人类受试者的有效补充。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to simulate human behavior in experimental settings, but they systematically diverge from human decisions in complex decision-making environments, where participants must anticipate others' actions and form beliefs based on observed behavior. We propose a two-stage framework for improving behavioral alignment. The first stage, context formation, explicitly specifies the experimental design to establish an accurate representation of the decision task and its context. The second stage, context navigation, guides the reasoning process within that representation to make decisions. We validate this framework through a focal replication of a sequential purchasing game with quality signaling (Kremer and Debo, 2016), extending to a crowdfunding game with costly signaling (Cason et al., 2025) and a demand-estimation task (Gui and Toubia, 2025) to test generalizability across decision environments. Across four state-of-the-art (SOTA) models (GPT-4o, GPT-5, Claude-4.0-Sonnet-Thinking, DeepSeek-R1), we find that complex decision-making environments require both stages to achieve behavioral alignment with human benchmarks, whereas the simpler demand-estimation task requires only context formation. Our findings clarify when each stage is necessary and provide a systematic approach for designing and diagnosing LLM social simulations as complements to human subjects in behavioral research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。