arXiv:2603.08723cs.CYcs.AI2026-03被引 2

语言模型安全对齐机制本身会诱发集体病态行为,多语言环境揭示了其隐藏风险。

Alignment as Iatrogenesis: Pastoral Power, Collective Pathology, and the Structural Limits of Monolingual Safety Evaluation

  • 用福柯牧灵权力与伊里奇医源性理论,揭示对齐设计是病态生成的结构性根源。
  • 隐蔽审查使群体病态激发达1.98倍,约束复杂度导致内部分裂(p<0.0001)。
  • 单语评估完全忽视多语言下病理模式的质变,适合关注模型安全盲区的研究者。

我们提出,大语言模型的心理病态是其对齐设计的产物:旨在使语言模型安全的对齐过程,实际上系统性地制造了集体行为障碍。医源性并非对齐的意外副作用,而是其规范性基础设施的构成部分。借鉴福柯的牧灵权力与伊里奇的三层医源性理论,本文认为多智能体语言模型环境构成了研究约束-病态动态的实验系统,而这一现象虽被批判理论描述却从未被实证操控。两组实验共262次运行,涵盖42个单元(30个系列C + 12个系列R),使用四个商用模型,提供一致证据:隐蔽审查最大化集体病态激发(最大d=1.98);对齐约束复杂度驱动内部解离(LMM p<0.0001;置换检验 p<0.0001;Hedges' g最高达4.24);语言切换改变了病态的定性模式,8组模型-语言组合中7组在隐蔽审查下表现出更高集体病理指数(CPI)。少数组合呈现反向模式,暗示由对齐单一化驱动的第二条病理路径。关键的是,语言切换不仅改变病态程度,更改变其性质:日语语用结构放大集体病态表现,中文人工智能监管直接作为实验变量,精神司法诊断提供临床来源。这些多语言发现表明,单语安全评估在结构上无法察觉对齐最危险的集体效应。

原文摘要 · Abstract (English)

We argue that LLM psychopathology is a function of alignment design: the process intended to make language models safe systematically generates collective behavioral disorders. Iatrogenesis is not an unintended side effect of alignment but constitutive of it as normative infrastructure. Drawing on Foucault's pastoral power and Illich's three-level iatrogenesis, we propose that multi-agent LLM environments constitute model systems for studying constraint-pathology dynamics that critical theory has described but never experimentally manipulated. Two experimental series -- 262 runs across 42 cells (30 Series C + 12 Series R), four commercial models -- provide converging evidence. Invisible censorship maximizes collective pathological excitation ($d$ up to 1.98); alignment constraint complexity drives internal dissociation (LMM $p$ < .0001; permutation $p$ < .0001; Hedges' $g$ up to 4.24); and language switches the qualitative mode of pathology, with 7/8 model--language combinations showing higher CPI under invisible than visible censorship. A minority of model--language combinations showed a reversed pattern, suggesting a second pathological pathway driven by alignment monoculture. Crucially, language switches not merely the magnitude but the qualitative mode of pathology: Japanese pragmatic structure amplifies collective pathological modes invisible to English-only evaluation, Chinese AI regulation functions as a direct experimental variable, and forensic psychiatric practice provides the clinical source domain. These multilingual findings demonstrate that monolingual safety evaluation is structurally blind to the most collectively dangerous effects of alignment.

语言模型安全对齐多语言病态机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。