arXiv:2504.14287cs.CLcs.CY2025-04

研究大模型如何被细微意识形态操控,发现微调比提示词影响更大。

Probing the Subtle Ideological Manipulation of Large Language Models

  • 构建多任务数据集,覆盖从进步左翼到保守右翼的完整意识形态谱系。
  • 微调后模型在政治立场上显著更贴合目标意识形态,提示词效果有限。
  • 揭示模型易受隐性操控风险,适合关注AI伦理与安全的研究者参考。

大型语言模型(LLMs)已深刻改变自然语言处理,但其在政治敏感领域的意识形态易操纵性引发关注。以往研究集中于左右二元偏见,通过显式提示和政治问答数据集微调。本文突破这一框架,探索模型在从进步左翼到保守右翼的完整意识形态光谱中被影响的程度。我们提出一个新型多任务数据集,包含意识形态问答、陈述排序、宣言完形填空及国会法案理解等任务,以反映多元立场。对Phi-2、Mistral和Llama-3三款模型在此数据集上进行微调,评估其采纳并表达复杂意识形态的能力。结果表明,微调显著增强模型的精细意识形态对齐,而显式提示仅带来微弱调整。这凸显了模型对微妙意识形态操控的高度敏感性,亟需建立更可靠的防护机制。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have transformed natural language processing, but concerns have emerged about their susceptibility to ideological manipulation, particularly in politically sensitive areas. Prior work has focused on binary Left-Right LLM biases, using explicit prompts and fine-tuning on political QA datasets. In this work, we move beyond this binary approach to explore the extent to which LLMs can be influenced across a spectrum of political ideologies, from Progressive-Left to Conservative-Right. We introduce a novel multi-task dataset designed to reflect diverse ideological positions through tasks such as ideological QA, statement ranking, manifesto cloze completion, and Congress bill comprehension. By fine-tuning three LLMs-Phi-2, Mistral, and Llama-3-on this dataset, we evaluate their capacity to adopt and express these nuanced ideologies. Our findings indicate that fine-tuning significantly enhances nuanced ideological alignment, while explicit prompts provide only minor refinements. This highlights the models' susceptibility to subtle ideological manipulation, suggesting a need for more robust safeguards to mitigate these risks.

大模型意识形态微调安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。