arXiv:2608.24232cs.AI2026-08

构建首个覆盖推理全过程的安全评估基准,揭示模型对中间推理步骤的防护短板。

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

论文配图:TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
图 1 · 摘自论文原文
  • 设计多语言、多风险类别的完整推理链评估体系
  • 发现推理过程安全判断难度远超输入与输出阶段
  • 适合研究大模型安全检测与可解释性的人群

大型推理模型(LRMs)在生成最终回答时看似安全,但其内部推理过程可能包含不安全内容。现有安全评测基准主要关注输入提示和最终输出,忽略对推理链条的评估,且仅提供二分类标签,缺乏判别依据。为此,我们提出TRACE——一个面向整个推理流程(提示、推理链条、最终输出)的证据驱动型安全评测基准。该基准涵盖双语提示,涉及九类风险与十种攻击策略;每条提示由四种LRM生成推理链与最终输出,并对各环节进行安全标注及证据提取。在TRACE上评估18个守门模型发现:推理链的安全判断显著更难,且当前模型难以准确提取支持性证据。结果凸显了对全推理链实现可靠且精准的安全检测的迫切需求。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchmarks for unsafe content detection focus primarily on prompts and final responses, leaving reasoning traces largely unexamined. Moreover, these benchmarks typically provide only binary safety labels, without evidence annotations that justify the judgments. To address these limitations, we introduce TRACE, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses. TRACE includes prompts in two languages spanning nine risk categories and ten attack strategies. For each prompt, four LRMs generate reasoning traces and final responses, and we annotate the safety of each component and extract supporting evidence from the corresponding source text. Evaluating 18 guardrail models on TRACE reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models struggle to accurately extract supporting evidence. These findings highlight the need for guardrail models that can reliably detect and precisely localize unsafe content across the LRM inference pipeline.

大模型安全推理链评测基准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。