arXiv:2508.08846cs.CLcs.AI2025-08中稿 · CASE@RANLP2025被引 7

通过分析模型内部表示,识别并修正大模型的政治偏见。

Steering Towards Fairness: Mitigating Political Bias in LLMs

  • 基于政治光谱测试,用对比样本提取模型隐藏层激活
  • 发现解码器模型各层均存在系统性政治表征偏差
  • 提出可迁移的向量修正方法,适合安全可控的AI应用

大型语言模型(LLMs)在现实应用中日益普及,但其在政治与经济维度上存在编码和传播意识形态偏见的问题。本文基于政治光谱测试(Political Compass Test, PCT),采用对比对分析解码器架构模型(如Mistral、DeepSeek)的内部表示,构建了跨层次、多意识形态轴的激活提取流程。结果表明,解码器模型在不同层中系统性地编码政治表征偏差,且这些偏差可被用于基于引导向量的有效干预。该研究揭示了政治偏见在模型中的深层编码机制,提供了一种超越表面输出修正的系统性去偏方法。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have enabled their widespread use across diverse real-world applications. However, concerns remain about their tendency to encode and reproduce ideological biases along political and economic dimensions. In this paper, we employ a framework for probing and mitigating such biases in decoder-based LLMs through analysis of internal model representations. Grounded in the Political Compass Test (PCT), this method uses contrastive pairs to extract and compare hidden layer activations from models like Mistral and DeepSeek. We introduce a comprehensive activation extraction pipeline capable of layer-wise analysis across multiple ideological axes, revealing meaningful disparities linked to political framing. Our results show that decoder LLMs systematically encode representational bias across layers, which can be leveraged for effective steering vector-based mitigation. This work provides new insights into how political bias is encoded in LLMs and offers a principled approach to debiasing beyond surface-level output interventions.

大模型偏见去偏方法模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。