arXiv:2604.24429cs.CL2026-04

评估政治对齐大模型的多维表现,揭示其潜在风险与权衡。

A Multi-Dimensional Audit of Politically Aligned Large Language Models

论文配图:A Multi-Dimensional Audit of Politically Aligned Large Language Models
图 1 · 摘自论文原文
  • 基于哈贝马斯沟通理论,从四维度量化评估模型
  • 大模型更真实但偏见更多,微调比角色扮演更公平
  • 适合关注AI伦理与政治应用的研究者与从业者

随着大语言模型(LLMs)在各行业广泛应用,其在政治话语等敏感领域被滥用的风险日益突出。通过提示工程或微调技术刻意对齐特定政治意识形态虽有助于政治宣传等场景,但也可能引发性能下降、误导信息传播及偏见加剧等问题。本文提出一种受哈贝马斯沟通行动理论启发的多维审计框架,从有效性、公平性、真实性与说服力四个维度,采用自动化定量指标评估九个经微调或角色扮演方式对齐的政治化LLMs。结果表明:更大模型在意识形态角色扮演和回答真实性上表现更好,但公平性更低,对不同意识形态者表现出更高水平的愤怒与攻击性语言;微调模型相比角色扮演模型偏见更低、对齐更有效,但在推理任务中表现下降且幻觉增多。所有测试模型均至少在一个维度存在缺陷,凸显了需发展更平衡、鲁棒的对齐策略。本研究旨在确保政治对齐模型生成合法、无害的论述,为负责任的政治对齐提供评估框架。

原文摘要 · Abstract (English)

As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse, especially in sensitive areas such as political discourse. Deliberately aligning LLMs with specific political ideologies, through prompt engineering or fine-tuning techniques, can be advantageous in use cases such as political campaigns, but requires careful consideration due to heightened risks of performance degradation, misinformation, or increased biased behavior. In this work, we propose a multi-dimensional framework inspired by Habermas' Theory of Communicative Action to audit politically aligned language models across four dimensions: effectiveness, fairness, truthfulness, and persuasiveness using automated, quantitative metrics. Applying this to nine popular LLMs aligned via fine-tuning or role-playing revealed consistent trade-offs: while larger models tend to be more effective at role-playing political ideologies and truthful in their responses, they were also less fair, exhibiting higher levels of bias in the form of angry and toxic language towards people of different ideologies. Fine-tuned models exhibited lower bias and more effective alignment than the corresponding role-playing models, but also saw a decline in performance reasoning tasks and an increase in hallucinations. Overall, all of the models tested exhibited some deficiency in at least one of the four metrics, highlighting the need for more balanced and robust alignment strategies. Ultimately, this work aims to ensure politically-aligned LLMs generate legitimate, harmless arguments, offering a framework to evaluate the responsible political alignment of these models.

大模型对齐政治偏见模型评估伦理风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。