arXiv:2501.14294cs.CLcs.AI2025-01ICLR被引 3

发现大模型易受政治刻板印象影响,放大党派立场偏差。

Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes

  • 用认知科学中的代表性启发法分析模型偏见
  • 模型比人类更夸大党派立场,加剧刻板印象
  • 提示词干预可有效缓解模型偏见

大语言模型(LLMs)在政治领域与人类价值观的对齐问题日益重要。已有研究显示,模型生成内容可能包含政治倾向,并模仿政党在议题上的立场。然而,模型偏离实际立场的程度及条件尚不明确。本研究基于认知科学中的代表性启发法(representativeness heuristics),即人们倾向于依据典型特征过度推断群体属性,分析模型如何产生政治刻板印象。通过实验发现,尽管模型能模仿某些政党的立场,但其夸大程度超过人类调查受访者;且模型对代表性的依赖远高于人类。这表明大模型易受代表性启发法影响,存在加剧政治刻板印象的风险。同时测试了提示词缓解策略,发现能减轻人类中代表性偏差的提示方法同样可降低模型响应中的代表性影响。

原文摘要 · Abstract (English)

Examining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them. Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses.

大模型对齐政治偏见认知启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。