让大模型模拟不同社会身份,解决意见仿真中同质化问题
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
- 在隐藏层注入参数化社会身份特征,实现细粒度控制
- 使仿真结果与真实调查数据的分布差异降低40%以上
- 适合做社会学研究、政策模拟的学者和机构使用
大语言模型(LLMs)被用于公共意见模拟,可替代成本高、耗时长的人类调查。然而现有方法难以捕捉社会多样性,导致群体间差异模糊、响应趋于同质。我们发现这是由于LLM隐层表示中出现“多样性坍缩”现象,不同社会身份逐渐不可区分。为此提出参数化社会身份注入(PSII)框架,将人口属性与价值取向等显式参数直接注入到LLM中间隐藏状态。相比提示词式人格设定,PSII可在表示层面实现更精细、可控的身份调节。在世界价值观调查(World Values Survey)数据集上,使用多个开源LLM进行实验表明,PSII显著提升分布保真度与多样性,使与真实数据的KL散度下降超40%,整体多样性明显增强。本工作为控制LLM代理的表示层级提供了新思路,推动可扩展、具多样性的公共意见仿真发展。
原文摘要 · Abstract (English)
Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability, current LLM-based simulation methods fail to capture social diversity, producing flattened inter-group differences and overly homogeneous responses across demographic groups. We identify this limitation as a Diversity Collapse phenomenon in LLM hidden representations, where distinct social identities become increasingly indistinguishable across layers. Motivated by this observation, we propose Parametric Social Identity Injection (PSII), a general framework that injects explicit, parametric representations of demographic attributes and value orientations directly into intermediate hidden states of LLMs. Unlike prompt-based persona conditioning, PSII enables fine-grained and controllable identity modulation at the representation level. Extensive experiments on the World Values Survey using multiple open-source LLMs show that PSII significantly improves distributional fidelity and diversity, reducing KL divergence to real-world survey data while enhancing overall diversity. This work provides new insights into representation-level control of LLM agents and advances scalable, diversity-aware public opinion simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。