用角色设定提升大模型安全,效果远超传统方法
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
- 用社会角色隐式编码价值观和推理模式
- 使深求-V3在野路子攻击下不安全输出从81.4%降至3.6%
- 无需训练,适合构建可解释的AI评判系统
基于心智理论,我们提出角色条件化作为原则对齐的简洁替代方案:社会角色(如母亲、法官)隐式包含价值与应用这些价值所需的认知框架。提出无训练流程,包含角色条件生成器与迭代角色批判器进行优化。在五个模型家族中,该方法在多个基准测试上持续优于基于原则、思维链(CoT)及其他基线。特别地,使用DeepSeek-V3时,野路子攻击(WildJailbreak)下的不安全输出率从81.4%降至3.6%。不仅适用于通用安全评测,还能稳定应用于智能体安全任务。结果确立角色分配作为强大且可解释的AI对齐范式,以及大语言模型作为裁判的构建方式。
原文摘要 · Abstract (English)
Principle-based alignment often lacks context sensitivity and completeness. Grounded in Theory of Mind, we propose role conditioning as a compact alternative: social roles (e.g., mother, judge) implicitly encode both values and the cognitive schemas required to apply them. We introduce a training-free pipeline featuring a role-conditioned generator and iterative role-based critics for refinement. Across five model families, our approach consistently outperforms principle-based, Chain-of-Thought (CoT) and other baselines across benchmarks. Notably, it reduces unsafe outputs on the WildJailbreak benchmark from 81.4\% to 3.6\% with DeepSeek-V3. Not only for common safety benchmarks, it consistently applies for agentic safety tasks. These results establish role assignment as a powerful, interpretable paradigm for AI alignment and LLM-as-a-Judge construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。