arXiv:2506.04774cs.CLcs.AI2025-06AAAI被引 2

用四维向量解析大模型内部政治立场,实现精准干预与检测

Fine-Grained Interpretation of Political Opinions in Large Language Models

  • 构建四维政治概念框架,通过可解释技术学习细粒度向量表征
  • 在8个开源模型上验证,向量能有效解耦政治概念混淆,泛化能力强
  • 适合研究模型偏见、安全对齐或内容调控的AI从业者使用

现有大模型政治立场研究多依赖开放生成响应,但存在输出与内部意图不一致的问题。本文提出从内部机制入手,突破单维度分析局限,设计四维政治学习框架并构建对应数据集,实现细粒度政治概念向量学习。采用三种可解释表示工程方法,在8个开源大模型上验证,结果表明该向量能有效解耦政治概念混淆,检测任务证实其语义合理性,并在分布外(OOD)设置下表现出良好泛化与鲁棒性。干预实验进一步证明,通过操控这些向量可引导模型生成不同政治倾向的回复。

原文摘要 · Abstract (English)

Studies of LLMs' political opinions mainly rely on evaluations of their open-ended responses. Recent work indicates that there is a misalignment between LLMs' responses and their internal intentions. This motivates us to probe LLMs' internal mechanisms and help uncover their internal political states. Additionally, we found that the analysis of LLMs' political opinions often relies on single-axis concepts, which can lead to concept confounds. In this work, we extend the single-axis to multi-dimensions and apply interpretable representation engineering techniques for more transparent LLM political concept learning. Specifically, we designed a four-dimensional political learning framework and constructed a corresponding dataset for fine-grained political concept vector learning. These vectors can be used to detect and intervene in LLM internals. Experiments are conducted on eight open-source LLMs with three representation engineering techniques. Results show these vectors can disentangle political concept confounds. Detection tasks validate the semantic meaning of the vectors and show good generalization and robustness in OOD settings. Intervention Experiments show these vectors can intervene in LLMs to generate responses with different political leanings.

大模型对齐可解释性政治立场向量干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。