arXiv:2508.21448cs.CL2025-08

模型拒绝回答政治指令,可能是因为能力不足而非安全限制。

When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

  • 用指令引导和激活控制方法测试模型的政治可引导性
  • 高可引导模型激活约7.3倍更多政治特征,低可引导模型则频繁拒绝
  • 删减少量政治特征会引发拒绝行为,证明能力缺陷是主因

大型语言模型有时会拒绝执行无害指令,如不支持某政治立场或扮演特定角色,这通常被视为安全机制在起作用。本文探讨这些拒绝是否反映的是模型内部表征能力的不足。为此提出「意识形态深度」概念,包含两个维度:(i)模型遵循政治指令的能力(可引导性),(ii)内部政治表征的特征丰富度,通过稀疏自编码器(SAEs)测量。基于两个主流开源大模型进行实验,比较提示与激活调控干预效果,并利用公开SAEs探测政治特征。结果显示显著差异:一个模型在双向政治指令下更具可引导性,其激活的政治特征数量高出约7.3倍;另一个模型则更多表现为拒绝。对前者进行小范围政治特征因果删除后,模型表现退化为特征贫乏状态并增加拒绝率。结果表明,某些拒绝行为源于能力缺陷而非固定安全规则,且意识形态深度是可量化的模型属性,有助于预测拒绝发生时机。

原文摘要 · Abstract (English)

Large language models (LLMs) sometimes refuse to follow benign instructions, such as declining to argue a political position or adopt a stated persona, and such refusals are commonly read as safety guardrails at work. We ask whether they can instead signal a **capability deficit**: a shortage of the internal representations a model needs to reason from the instructed perspective. To investigate, we introduce **ideological depth**, a property with two components: (i) a model's ability to follow political instructions without *failure* (steerability), and (ii) the **feature richness** of its internal political representations, measured with sparse autoencoders (SAEs). Using two widely used openweight LLMs as candidates, we compare interventions based on prompts and activation-steering, and probe political features with publicly available SAEs. We find large, systematic differences: a model that is more steerable in both ideological directions activates **~7.3x** more distinct political features, while the other model instead responds with increased refusals. Causally ablating a small, targeted set of political features from the former model reproduces the same feature-poor behavior and drives up refusals. Together, these results indicate that refusals on benign prompts can arise from **capability deficits** rather than fixed safety rules, and that ideological depth is a measurable property of LLMs that helps predict when a model will refuse.

大模型政治推理特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。