arXiv:2607.23519cs.CYcs.AI2026-07中稿 · AIES 2026

测试大模型在政治轴上的可调控性,发现指令影响远超模型自身立场。

Auditing Alignment Controllability in LLMs via Political Axes

论文配图:Auditing Alignment Controllability in LLMs via Political Axes
图 1 · 摘自论文原文
  • 通过12种意识形态角色和70个政治议题,测试指令对模型输出的操控效果。
  • 上下文引导解释了88%-93%的政治经济与社会维度差异,模型身份影响不足3%。
  • 建议引入漂移、饱和、拒绝阈值等指标,评估模型真实可控性,适合安全与伦理研究者。

大语言模型的政治审计常将其简化为政治光谱上的一点,但实际部署中更重要的是模型回答可被引导的程度与方向。这种引导依赖系统提示,即平台设定或用户历史生成的个性化层,而非人工手写。我们针对7个主流大模型(GPT-5、Claude、Grok、Gemini、DeepSeek、Kimi、Qwen)开展以分散性为核心的可控性压力测试,涵盖12种意识形态人格及一个无引导基线,覆盖70个政治光谱问题,每项重复10次,共收集63,700条响应。结果显示,语境框架可解释经济与社会轴上88%-93%的方差,模型身份影响小于3%,表明输出高度依赖指令调整。不同模型响应模式各异:部分更易移动,部分在极端提示下趋于饱和。先前审计中矛盾的定向引导结果,在识别基线非中心化后得以统一:位移与接近度分化,效应呈几何特征,非单纯合规差异。在威权提示下,模型对相同问题产生相似偏移。因此,政治坐标审计应辅以报告分散性、对称性、饱和度与拒绝阈值的可调控性审计。本文公开所有提示、基准数据与代码。

原文摘要 · Abstract (English)

Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.

模型可控性政治偏见大模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。