用可解释的因果树评估AI行为风险,支持灵活调整安全标准。
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
- 构建危害-收益树分析AI行为的潜在影响,标注可能性、严重性和紧迫性。
- 在19,000个提示上训练,模型在分类任务中平均F1达0.81,优于现有系统。
- 参数可调,适合需要透明安全策略的团队或社区使用。
理想的AI安全审核系统应具备结构可解释性(便于可靠解释决策)和可调节性(以对齐安全标准并反映社区价值观),但现有系统尚不满足。为此,我们提出SafetyAnalyst——一种新型AI安全审核框架。给定一个AI行为,SafetyAnalyst通过链式思维推理生成结构化的“危害-收益树”,列举该行为可能引发的有害与有益行动及后果,并标注其可能性、严重性和紧迫性,评估对利益相关方的影响。随后,利用28个完全可解释的权重参数,将所有影响聚合为一个可解释的危害评分,该评分可按特定安全偏好进行调整。我们基于此框架开发了一个开源的大型语言模型提示安全分类系统,该系统由前沿大模型在1.9万条提示上生成的1850万个危害-收益特征提炼而来。在综合基准测试中,SafetyAnalyst(平均F1=0.81)在提示安全分类任务上优于现有系统(平均F1<0.72),同时具备可解释性、透明性和可调节性优势。
原文摘要 · Abstract (English)
The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on. To address this gap, we present SafetyAnalyst, a novel AI safety moderation framework. Given an AI behavior, SafetyAnalyst uses chain-of-thought reasoning to analyze its potential consequences by creating a structured "harm-benefit tree," which enumerates harmful and beneficial actions and effects the AI behavior may lead to, along with likelihood, severity, and immediacy labels that describe potential impacts on stakeholders. SafetyAnalyst then aggregates all effects into a harmfulness score using 28 fully interpretable weight parameters, which can be aligned to particular safety preferences. We applied this framework to develop an open-source LLM prompt safety classification system, distilled from 18.5 million harm-benefit features generated by frontier LLMs on 19k prompts. On comprehensive benchmarks, we show that SafetyAnalyst (average F1=0.81) outperforms existing moderation systems (average F1$<$0.72) on prompt safety classification, while offering the additional advantages of interpretability, transparency, and steerability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。