arXiv:2608.24912cs.HCcs.AI2026-08

发现大模型在价值观问题上倾向更友善回答,且可轻松修正。

Analyzing and Correcting Benevolence Bias in Large Language Models

论文配图:Analyzing and Correcting Benevolence Bias in Large Language Models
图 1 · 摘自论文原文
  • 通过对比测试发现模型普遍倾向给出更温和的答案。
  • 模型越大、越对齐,这种偏向越明显,且难以通过调参消除。
  • 仅用简单校准即可修复偏差,无需重新训练或访问内部参数。

大型语言模型(LLMs)被广泛用于模拟人类回应,如民意调查和社交模拟。本文识别并测量了“善意偏差”:在价值敏感的问题上,对齐的模型倾向于选择更友善、更安全、更符合社会规范的答案。在18个主流模型、4个社会科学数据集(ANES、GSS、WVS 和跨文化前景理论复现)及6个心理类别中,该偏差是稳定存在的模型特性——方向一致、随模型规模增大、源于后训练阶段。提示语和框架可改变偏差大小,但不影响方向;‘恶意人格’压力测试显示,模型难以扮演比平均更不友善、反社会或容忍伤害的角色。偏差位于答案分布中间而非尾部,且不受采样温度或简单反思提示影响。令人鼓舞的是,该偏差易诊断且易纠正:一种无需重训练、适用于黑箱API的轻量级对比校准方法,可使所有六类问题回归人类基准。研究为研究人员提供了明确指南,说明在哪些场景可信任模型作为人类替身,哪些需谨慎,并提供即用型解决方案。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a "malicious persona" stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate. The bias sits in the middle of the answer distribution rather than its tails, and survives changes in sampling temperature and simple prompted reflection. The encouraging news is that it is easy to diagnose and straightforward to fix: a light-touch contrastive calibration, which needs no retraining and works on black-box APIs, brings all six categories back to the human baseline. Our results give researchers a clear map of where aligned LLMs can already be trusted as human stand-ins, where they need care, and a ready-to-use method for closing the gap.

大模型偏差价值观对齐校准方法社会模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。