构建细粒度安全行为标注数据集,实现大模型推理中的安全行为实时检测与干预。
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
- 在句子级别标注推理过程中的安全行为,支持激活值层面的监控。
- 通过提取引导向量,可精准识别并干预模型中的危险推理行为。
- 适合研究大模型安全、提示攻击防御及可解释性方向的学者使用。
近期研究强调了监控大模型链式思维对AI安全的重要性;然而,现有基于文本的分析方法可能遗漏细微的有害模式,且易被隐藏不安全推理的模型绕过。本文提出一个句级标注的数据集,支持在大模型推理过程中基于激活值监测安全行为。该数据集包含带有安全行为标注的推理序列,如表达安全担忧或推测用户意图等,用于提取可检测和调控这些行为的引导向量。该数据集填补了安全研究的关键空白:现有数据集多为整体标注,而通过精确定位特定行为在推理链中的发生时刻,可显著提升引导向量在安全监控中的应用效果。我们展示了该数据集的实用性,成功提取出能同时检测与引导安全行为的激活表示,验证了激活层技术在推理安全监督中的潜力。内容警告:本文讨论有害提示情境下的AI安全问题,可能涉及潜在有害内容。
原文摘要 · Abstract (English)
Recent work has highlighted the importance of monitoring chain-of-thought reasoning for AI safety; however, current approaches that analyze textual reasoning steps can miss subtle harmful patterns and may be circumvented by models that hide unsafe reasoning. We present a sentence-level labeled dataset that enables activation-based monitoring of safety behaviors during LLM reasoning. Our dataset contains reasoning sequences with sentence-level annotations of safety behaviors such as expression of safety concerns or speculation on user intent, which we use to extract steering vectors for detecting and influencing these behaviors within model activations. The dataset fills a key gap in safety research: while existing datasets label reasoning holistically, effective application of steering vectors for safety monitoring could be improved by identifying precisely when specific behaviors occur within reasoning chains. We demonstrate the dataset's utility by extracting representations that both detect and steer safety behaviors in model activations, showcasing the potential of activation-level techniques for improving safety oversight on reasoning. Content Warning: This paper discusses AI safety in the context of harmful prompts and may contain references to potentially harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。