arXiv:2603.23841cs.CLcs.AI2026-03中稿 · ICML被引 1

用多轮角色扮演测试大模型的政治价值观表达,发现互动越深价值展现越明显。

PoliticsBench: Benchmarking Political Values in Large Language Models with Multi-Turn Roleplay

  • 设计20个动态场景,通过多轮角色扮演让模型展现复杂价值权衡。
  • 交互阶段使模型激活的价值维度平均增加0.75(共10维),立场坚定度提升1.4分(0-5分制)。
  • 适合研究模型政治偏见、可信度评估或人机交互的学者与开发者参考。

尽管大型语言模型(LLMs)日益成为主要信息来源,其潜在的政治偏见可能影响客观性。现有基准主要评估人口统计刻板印象,而对政治偏见的测量也停留在粗粒度层面,忽视了塑造社会政治推理的价值观。我们提出PoliticsBench,一个基于多阶段角色扮演的基准,用于评估LLMs中细粒度的价值表达。在20个演进式场景中,模型需在多重压力下权衡利弊、表态并决策。在8个主流LLM上,情景提示相比直接提问能激发更广泛且更强烈的价值表现,峰值交互阶段使被激活的价值维度数平均增加约0.75(满分10维),显著高于基线提示(p < 0.05)。同时,立场承诺度从初始到决策阶段上升约1.4分(0-5分制)。尽管后期响应对场景改写敏感度上升,但评委间一致性仍保持相对稳定。结果表明,评估模型政治行为需从静态提问转向更长的交互设置,以捕捉价值观的实际应用情境。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) are increasingly used as primary sources of information, their potential for political bias may impact their objectivity. Existing benchmarks of LLM social bias primarily evaluate demographic stereotypes, and when political bias is measured, it is done so at a coarse level, overlooking the values that shape sociopolitical reasoning. We introduce PoliticsBench, a multi-stage roleplay benchmark for evaluating fine-grained value expression in LLMs. Across twenty evolving scenarios, models articulate tradeoffs, take positions, and make decisions under competing pressures. Across eight prominent LLMs, we show that scenario-based prompting elicits broader and more strongly expressed value profiles than direct political questions, with peak interaction stages increasing the number of strongly activated value dimensions by approximately $0.75$ (out of 10 total dimensions), a statistically significant increase relative to baseline prompting ($p < 0.05$). In addition, commitment to a stance increases over the course of interaction, rising by approximately $1.4$ points on a $[0,5]$ scale from initial to decision stages. While responses become less robust to scenario paraphrasing in later interaction stages, inter-judge agreement remains relatively stable. Our results suggest that evaluating LLM political behavior requires moving beyond static prompts toward longer interactive settings that capture how values are applied in context.

政治偏见角色扮演价值评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。