arXiv:2512.22725cs.CLcs.CY2025-12被引 2

用简单提示词调整,让大模型回答更贴近真实人群。

Mitigating Social Desirability Bias in Random Silicon Sampling

  • 用中性第三人称重写问题,减少社会期望偏差。
  • 重写提示使模型输出分布与真实数据差距缩小30%以上。
  • 适合做民意模拟的研究者参考实用技巧。

大语言模型(LLMs)被用于模拟人口回应,称为“硅基抽样”。然而,在涉及社会敏感问题时,模型常表现出社会期望偏差,答案偏向符合社会规范。本文基于美国全国选举研究(ANES)数据,测试了三种模型(Llama-3.1系列和GPT-4.1-mini)在四种提示策略下的表现:重写提示(中性第三人称)、反向编码(语义反转)、引导提示和前导指令,分别鼓励分析性和真诚性。通过詹森-申农散度评估与真实数据的对齐程度。结果表明,重写提示最有效,显著降低对社会可接受答案的集中倾向,使分布更接近真实数据;反向编码效果不一;引导和前导指令导致回答趋于均匀,但未系统改善偏差。研究验证了提示框架控制在缓解模型固有社会期望偏差中的有效性,为构建更具代表性的硅基样本提供了可行路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social desirability bias in LLM-based sampling remain limited. In this work, we investigate whether minimal, psychologically grounded prompt wording can mitigate this bias and improve alignment between silicon and human samples. We conducted a study using data from the American National Election Study (ANES) on three LLMs from two model families: the open-source Llama-3.1 series and GPT-4.1-mini. We first replicate a baseline silicon sampling study, confirming the persistent Social Desirability Bias. We then test four prompt-based mitigation methods: \emph{reformulated} (neutral, third-person phrasing), \emph{reverse-coded} (semantic inversion), and two meta-instructions, \emph{priming} and \emph{preamble}, respectively encouraging analytics and sincerity. Alignment with ANES is evaluated using Jensen-Shannon Divergence with bootstrap confidence intervals. Our results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. Reverse-coding produced mixed results across eligible items, while the Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation. Our findings validate the efficacy of prompt-based framing controls in mitigating inherent Social Desirability Bias in LLMs, providing a practical path toward more representative silicon samples.

大模型社会偏差提示工程民意模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。