解决安全对齐中模型表达力下降问题,提升大模型安全性与有用性。
Constitutional On-Policy Safe Distillation

- 用宪法条件引导教师模型,防止输出过度保守
- 在3个多模态大模型上,安全与有用性双提升
- 适合需要高安全性和强表达能力的场景
基于特权信息的教师模型提供密集的词元级监督,使在线策略自蒸馏(OPSD)成为高效的后训练范式。然而,现有研究表明,在可验证推理任务中OPSD会严重坍缩;而安全对齐依赖高层宪法而非明确答案,仍存在类似问题:宪法约束导致教师分布收缩为短且过于保守的响应,反KL进一步加剧表达力下降。我们将其归因于非正交语义空间中的几何泄漏,安全压力渗入表达维度。为此提出宪法在线策略安全蒸馏(COPSD),先通过跨SFT冷启动校准教师,再进行宪法条件下的在线蒸馏。在3个多模态大语言模型上,12项安全与通用基准测试显示,相比基线,COPSD同时提升安全性与帮助性,并降低对通用推理的安全代价。
原文摘要 · Abstract (English)
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervision. Prior work has shown that OPSD can collapse in verifiable reasoning tasks, while safety alignment differs in that it is guided by high-level constitutions rather than explicit target answers. However, pilot studies reveals that safety OPSD nonetheless suffers from severe collapse, where constitutional conditioning contracts the teacher distribution toward short and overly conservative responses and Reverse KL further amplifies this contraction into reduced expressiveness. We formalize this effect as geometric leakage under safety boundaries in a non-orthogonal semantic space, where safety pressure transfers into the expressiveness dimension. Based on this analysis, we propose Constitutional On-Policy Safety Distillation (COPSD), which first calibrates the teacher through a Cross-SFT cold-start and then performs constitution-conditioned on-policy distillation. Experiments across 3 multimodal large language models (MLLMs) on 12 safety and general benchmarks show that COPSD improves both safety and helpfulness over baselines while reducing the safety tax on general reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。