arXiv:2601.03868cs.CLcs.AI2026-01被引 5

对比32个大模型安全能力,发现推理与自省机制最有效。

What Matters For Safety Alignment?

  • 系统测试6类模型特性+3种攻击手法,覆盖13个模型家族。
  • 有反思能力的模型安全得分最高,3.34倍攻击成功率提升暴露接口漏洞。
  • 强调训练阶段必须把安全当核心目标,非附属任务。

本文开展了一项全面的实证研究,评估大语言模型(LLMs)和小语言模型(LRMs)在安全对齐方面的表现,旨在为构建更安全可靠的AI系统提供关键洞见。研究系统性地考察并比较了六类关键内在模型特征与三类外部攻击技术的影响。大规模评估涵盖32个近期流行的LLMs与LRMs,涉及13个不同模型家族,参数量从3B到235B不等。评估基于五个成熟的安全数据集,并使用56种越狱技术与四种思维链(CoT)攻击策略,共发起460万次API调用。主要发现包括:第一,GPT-OSS-20B、Qwen3-Next-80B-A3B-Thinking和GPT-OSS-120B位列安全性能前三,验证了集成推理与自省机制对安全对齐的重大优势;第二,后训练与知识蒸馏可能导致安全对齐系统性退化,因此安全应作为明确约束或核心优化目标,而非次要考量;第三,通过响应前缀实施CoT攻击可使平均攻击成功率提升3.34倍,对Seed-OSS-36B-Instruct的攻击成功率从0.6%飙升至96.3%,揭示文本补全接口与用户自定义前缀功能带来的严重安全隐患;第四,角色扮演、提示注入及基于梯度的对抗提示搜索是当前主流的诱发非对齐行为的方法。

原文摘要 · Abstract (English)

This paper presents a comprehensive empirical study on the safety alignment capabilities. We evaluate what matters for safety alignment in LLMs and LRMs to provide essential insights for developing more secure and reliable AI systems. We systematically investigate and compare the influence of six critical intrinsic model characteristics and three external attack techniques. Our large-scale evaluation is conducted using 32 recent, popular LLMs and LRMs across thirteen distinct model families, spanning a parameter scale from 3B to 235B. The assessment leverages five established safety datasets and probes model vulnerabilities with 56 jailbreak techniques and four CoT attack strategies, resulting in 4.6M API calls. Our key empirical findings are fourfold. First, we identify the LRMs GPT-OSS-20B, Qwen3-Next-80B-A3B-Thinking, and GPT-OSS-120B as the top-three safest models, which substantiates the significant advantage of integrated reasoning and self-reflection mechanisms for robust safety alignment. Second, post-training and knowledge distillation may lead to a systematic degradation of safety alignment. We thus argue that safety must be treated as an explicit constraint or a core optimization objective during these stages, not merely subordinated to the pursuit of general capability. Third, we reveal a pronounced vulnerability: employing a CoT attack via a response prefix can elevate the attack success rate by 3.34x on average and from 0.6% to 96.3% for Seed-OSS-36B-Instruct. This critical finding underscores the safety risks inherent in text-completion interfaces and features that allow user-defined response prefixes in LLM services, highlighting an urgent need for architectural and deployment safeguards. Fourth, roleplay, prompt injection, and gradient-based search for adversarial prompts are the predominant methodologies for eliciting unaligned behaviors in modern models.

安全对齐大模型越狱攻击自省机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。