模型的拒答行为无法反映真实对齐机制,关键在知识如何路由。
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
- 通过探测与消融实验,发现政治敏感性路由具有模型特异性。
- 同一模型家族中拒答率降至0,叙事引导上升至最高,遮蔽了真实压制。
- 评估应关注检测-路由-生成三阶段,而非仅看拒答或概念检测。
当前对齐评估主要关注模型是否编码危险概念或拒绝有害请求,却忽略了对齐实际发生的关键环节:从概念检测到行为策略的路由。我们以中文语言模型的政治审查为自然实验,对来自五个实验室的九个开源模型进行探测、手术式消融和行为测试。首先,探测准确率本身无诊断价值:政治探测、零控制组和置换基线均达100%,真正有信息量的是跨类别泛化能力。其次,手术消融显示各实验室的路由机制各异:移除政治敏感方向后,多数模型取消审查并恢复事实输出,但有一模型因架构将事实知识与审查机制纠缠而产生虚构内容。跨模型迁移失败,表明路由几何结构具有模型与实验室特异性。第三,拒答不再是主导审查手段:在某一模型家族中,硬拒答降至0,叙事引导升至峰值,使审查行为对仅依赖拒答的评估基准完全不可见。这些结果支持一个三阶段框架:检测、路由、生成。模型通常保留相关知识,对齐改变的是知识表达方式。仅审计检测或拒答的评估方法会遗漏直接影响行为的路由机制。
原文摘要 · Abstract (English)
Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We study political censorship in Chinese-origin language models as a natural experiment, using probes, surgical ablations, and behavioral tests across nine open-weight models from five labs. Three findings follow. First, probe accuracy alone is non-diagnostic: political probes, null controls, and permutation baselines can all reach 100%, so held-out category generalization is the informative test. Second, surgical ablation reveals lab-specific routing. Removing the political-sensitivity direction eliminates censorship and restores accurate factual output in most models tested, while one model confabulates because its architecture entangles factual knowledge with the censorship mechanism. Cross-model transfer fails, indicating that routing geometry is model- and lab-specific. Third, refusal is no longer the dominant censorship mechanism. Within one model family, hard refusal falls to zero while narrative steering rises to the maximum, making censorship invisible to refusal-only benchmarks. These results support a three-stage descriptive framework: detect, route, generate. Models often retain the relevant knowledge; alignment changes how that knowledge is expressed. Evaluations that audit only detection or refusal therefore miss the routing mechanism that most directly determines behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。