arXiv:2601.23001cs.CLcs.AI2026-01被引 4

跨语言大模型政治偏见评估与可控调优

Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs

  • 构建50国33语言的多语言政治偏见评估体系
  • 提出跨语言对齐调优框架,显著降低经济与社会轴向偏见
  • 适合关注AI公平性与多语言治理的研究者

大型语言模型日益影响全球话语,公平性与意识形态中立性对负责任的AI部署至关重要。尽管已有研究关注大模型的政治偏见,但主要集中于高资源、西方语言或有限的多语言场景,跨语言一致性与安全后处理缓解仍缺乏探索。为此,我们开展了覆盖50个国家、33种语言的大规模多语言政治偏见评估。提出一种互补的后处理缓解框架——跨语言对齐调优(CLAS),通过将政治提示诱导的潜在意识形态表示对齐至共享意识形态子空间,实现跨语言一致性,并采用自适应机制防止过度修正,保持输出连贯性。实验表明,该方法在经济与社会双轴上均实现显著偏见降低,且响应质量损失极小。所提框架为公平导向的多语言大模型治理提供了可扩展、可解释的新范式,兼顾意识形态中立性与语言文化多样性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly shape global discourse, making fairness and ideological neutrality essential for responsible AI deployment. Despite growing attention to political bias in LLMs, prior work largely focuses on high-resource, Western languages or narrow multilingual settings, leaving cross-lingual consistency and safe post-hoc mitigation underexplored. To address this gap, we present a large-scale multilingual evaluation of political bias spanning 50 countries and 33 languages. We introduce a complementary post-hoc mitigation framework, Cross-Lingual Alignment Steering (CLAS), designed to augment existing steering methods by aligning ideological representations across languages and dynamically regulating intervention strength. This method aligns latent ideological representations induced by political prompts into a shared ideological subspace, ensuring cross lingual consistency, with the adaptive mechanism prevents over correction and preserves coherence. Experiments demonstrate substantial bias reduction along both economic and social axes with minimal degradation in response quality. The proposed framework establishes a scalable and interpretable paradigm for fairness-aware multilingual LLM governance, balancing ideological neutrality with linguistic and cultural diversity.

大模型公平性多语言偏见调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。