arXiv:2607.11471cs.CL2026-07中稿 · Konvens 2026

测试大模型在复杂社会议题上的立场一致性,发现它们很少中立且常有惊人共识。

Are LLMs ready for HardChoices?

  • 构建新数据集HardChoices,评估模型对重大社会议题的立场
  • 无论大小模型,极少选择中立,多数表现出一致观点
  • 模型在分歧议题上反而更一致,揭示其潜在认知盲区

大量研究关注大语言模型(LLMs)是否存在政治偏见,但主要聚焦于左右或进步保守等高层意识形态维度。已有研究表明,尽管LLMs普遍呈现左倾和进步倾向,主要模仿训练数据中的偏见,但可通过后训练部分调整其偏好。本文通过一个新构建的数据集HardChoices,检验模型在重大实质性社会议题上的稳健立场。这些议题通常在同意识形态阵营内部也存在分歧。结果表明,面对此类问题时,无论是大型还是小型模型,均极少声明中立,常表现出不连贯性,并在表态的议题上展现出显著的一致性。

原文摘要 · Abstract (English)

A lot of research attention has been devoted to checking whether large language models (LLMs) are politically biased. This work has largely focused on high-level ideological dimensions, such as left--right or progressive--conservative, and it has been shown that while LLMs are predominantly left and progressive leaning, largely mimicking the biases in the training data, they can be to some extent steered to change their preferences in post-training. In this short note, we check if LLMs have robust stances with regard to major substantive societal issues, on which members of the same ideological camp are often in disagreement, summarised in a novel dataset \textsc{HardChoices}. We show that, faced with this line of questioning, LLMs, both large and small, surprisingly rarely declare neutrality, are often incoherent, and demonstrate a remarkable degree of agreement on issues where they do take stances.

大模型偏见社会议题立场一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。