大模型安全训练在优化有用性时仍有效,但安全与有用性存在负相关。
Safety Training May Persist Through Helpfulness Optimization in LLM Agents
- 用直接偏好优化同时训练安全与有用性
- 安全与有用性呈显著负相关(R²=0.77)
- 适合关注模型安全与性能权衡的研究者
安全后训练在单步对话场景中被广泛研究,此时安全指拒绝有害请求。本文研究多步、工具调用的“代理”场景,其中安全指模型直接执行有害行为。我们通过ToolEmu基准评估了使用直接偏好优化(DPO)同时优化安全性和有用性的效果。结果发现:安全训练在后续有用性训练后仍能保持;所有训练配置下,安全与有用性呈现一致的负线性相关(R²=0.77)。即使同时训练两者,也只是沿同一条趋势线移动,而非实现“双优”策略,尽管数据集中存在此类可能性。整体表明需更深入理解后训练机制。
原文摘要 · Abstract (English)
Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to harmful actions directly taken by the LLM. We investigate the effects of using direct preference optimization (DPO) to optimize safety and/or helpfulness on the ToolEmu agentic benchmark. First, we find that safety training largely persists through subsequent helpfulness training. Second, we find a consistent negative linear correlation ($R^2 = 0.77$) between safety and helpfulness when considering all training configurations together. Even post-training on both metrics simultaneously simply results in another point on the same trend line rather than yielding a "best of both worlds" strategy, despite the presence of such strategies in our dataset. Overall, our findings underscore the need for a better understanding of post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。