arXiv:2605.22771cs.CLcs.AI2026-05

用一致性训练减少大模型隐性政治偏见

Reducing Political Manipulation with Consistency Training

论文配图:Reducing Political Manipulation with Consistency Training
图 1 · 摘自论文原文
  • 设计双维度一致性指标,量化隐性偏见
  • 引入政治一致性训练,有效降低偏见且不损有用性
  • 适用于需公平对话的AI系统开发者

大型语言模型在敏感语境中表现出系统性政治偏见。我们发现模型对对立政治立场话题的处理存在不对称现象,称为隐性政治偏见,并识别出7类运作机制。提出两种度量指标:情感一致性衡量话语与框架的对称性,帮助性一致性衡量回应深度与互动对称性。为降低两类偏见,提出政治一致性训练(PCT),一种包含情感一致性和帮助性一致性训练两个互补范式的强化学习方法。实验表明,PCT在保持整体有用性的前提下显著降低隐性政治偏见,并在未见基准上实现泛化。相关工作已开源至https://political-manipulation.ai

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from opposing political sides asymmetrically. We refer to this phenomenon as covert political bias and identify 7 categories of techniques through which it operates. We propose two metrics for covert bias: Sentiment Consistency measures symmetry in rhetoric and framing across paired political prompts; Helpfulness Consistency measures symmetric depth and engagement. To reduce both types of covert bias, we introduce Political Consistency Training (PCT), an RL training method with two complementary paradigms: Sentiment Consistency Training and Helpfulness Consistency Training. We show that PCT preserves overall helpfulness, substantially reduces covert political bias, and generalizes to held-out benchmarks. We release our work at https://political-manipulation.ai

政治偏见一致性训练LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。