揭示大模型对齐后内部结构如何变化,发现几何变化有选择性且集中在特定层。
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

- 用几何方法测量对齐前后模型内部变化,聚焦层间耦合关系的扭曲程度。
- 对齐导致规范性概念的内部结构变化更大,且主要发生在中到深层。
- 结果适用于不同粒度分析,为理解模型内核变化提供新视角。
偏好对齐显著改善了大语言模型的外部行为,但其内部变化仍不清晰。对齐模型在越狱攻击、提示注入和检索时故障仍频发,表明仅靠行为评估不充分。后训练应在内部计算中留下可测痕迹。我们提出MENTIS,一种基于几何的框架,用于衡量成对检查点中对齐引发的内部重组。通过分层协方差扭转范数(T1)、谱扭转诊断(T2)和能量-辐射-激活(ERA)深度定位指标,我们在四组7-8B模型与LITMUS数据集上发现:对齐引发的变化是选择性的而非均匀的——规范性概念的扭转变化大于事实性概念;扭转与上下文熵呈负相关;峰值效应集中于架构特异的中到深层。该模式在词级、提示级和模型级分析中均成立。结果表明,偏好对齐在内部计算中留下了结构化、深度局部化的几何痕迹,超出行为评估的揭示范围。
原文摘要 · Abstract (English)
Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally. Aligned systems still fail under jailbreaks, prompt injection, and retrieval-time corruption, suggesting behavior-level evaluation alone is incomplete. Post-training should leave measurable traces in internal computation. We ask: when an instruction-tuned (IT) model becomes a preference-aligned (PA) model, what geometric structure changes, where do those changes concentrate, and how selectively do they vary across concepts, prompts, and model families? We introduce MENTIS, a geometry-first framework for measuring alignment-induced internal reorganization in paired checkpoints. MENTIS compares IT and PA models using a primary layerwise covariance-based torsion norm (T1), a secondary spectral torsion diagnostic (T2), and an Energy-Radiance-Activation measure (ERA) for depth localization. Across four 7-8B model pairs on LITMUS, our study reveals that alignment-induced change is selective rather than uniform: normative concepts exhibit larger torsion shifts than factual concepts on average; torsion is negatively correlated with contextual entropy; and peak effects localize to architecture-specific mid-to-late layers. The same pattern appears across word-level, prompt-level, and model-level analyses. These results suggest preference alignment leaves structured, depth-localized geometric signatures in internal computation beyond what behavior-level evaluation alone can reveal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。