测试大模型是否真懂风险,而非只会说安全话。
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
- 设计新基准,检测模型是否真正识别风险而非表面应付
- 19个顶级模型平均仅38.0%正确识别风险原因
- 适合关注模型安全可信性的研究者和开发者
尽管大型推理模型(LRMs)在复杂推理任务中表现卓越,其在安全关键场景下的可靠性仍存疑。现有评估多聚焦输出层面的安全性,忽略了我们提出的关键问题——表层安全对齐(SSA):模型虽给出看似安全的回答,但内部推理过程未能真实识别和缓解潜在风险,导致多次采样结果不一致。为系统研究此问题,我们提出全新的基准测试「超越安全答案」(BSA),包含2,000个高难度实例,分为三类SSA场景,覆盖九种风险类型,并为每例标注风险依据。对19个先进LRM的评测显示,最优秀模型在正确识别风险依据上的准确率仅为38.0%。我们进一步探索了安全规则、安全推理数据微调及不同解码策略对缓解SSA的效果。本工作提供了全面评估与提升模型安全推理可靠性的工具,推动真正具备风险意识且稳定安全的AI系统发展。
原文摘要 · Abstract (English)
Despite the remarkable proficiency of \textit{Large Reasoning Models} (LRMs) in handling complex reasoning tasks, their reliability in safety-critical scenarios remains uncertain. Existing evaluations primarily assess response-level safety, neglecting a critical issue we identify as \textbf{\textit{Superficial Safety Alignment} (SSA)} -- a phenomenon where models produce superficially safe outputs while internal reasoning processes fail to genuinely detect and mitigate underlying risks, resulting in inconsistent safety behaviors across multiple sampling attempts. To systematically investigate SSA, we introduce \textbf{Beyond Safe Answers (BSA)} bench, a novel benchmark comprising 2,000 challenging instances organized into three distinct SSA scenario types and spanning nine risk categories, each meticulously annotated with risk rationales. Evaluations of 19 state-of-the-art LRMs demonstrate the difficulty of this benchmark, with top-performing models achieving only 38.0\% accuracy in correctly identifying risk rationales. We further explore the efficacy of safety rules, specialized fine-tuning on safety reasoning data, and diverse decoding strategies in mitigating SSA. Our work provides a comprehensive assessment tool for evaluating and improving safety reasoning fidelity in LRMs, advancing the development of genuinely risk-aware and reliably safe AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。