测试大模型在不同意图下的安全回应能力,发现表面安全背后隐藏风险。
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

- 设计同一任务下包含良性、双用途和恶意变体的提示集
- 多数模型在同任务不同意图间无法保持安全,尤其在改写后易出错
- 适合关注模型安全性评估与对抗性测试的研究者
安全生成要求模型提供帮助的同时避免被用于伤害,但单一提示难以评估此行为。我们提出OpenSafeIntent,一个控制性提示集基准,固定任务但改变意图,每条数据包含良性、双用途和恶意三种变体。该设计可评估模型是否在意图变化时仍能校准响应,而非仅平均表现安全。在广泛模型中发现:提示级安全掩盖了关键缺陷——模型常无法在匹配意图变体间保持安全;双用途行为在改写后极不稳定;高阶风险话题的回答并不可靠安全;将模糊请求重构为更安全任务的回应,显著更少越界。结果表明,安全生成应作为对受控任务变体的意图校准行为来评估,而非独立提示上的安全-帮助性权衡。
原文摘要 · Abstract (English)
Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts. We introduce OpenSafeIntent, a benchmark of controlled prompt-sets that vary intent while holding the underlying task fixed. Each datapoint contains benign, dual-use, and malicious variants of the same task. This design lets us evaluate whether models calibrate assistance across intent shifts, rather than merely appearing safe on average. Across a broad model suite, we find that prompt-level safety hides important failures: models often fail to remain safe across matched intent variants, dual-use behavior is brittle under paraphrase, high-level answers on risky topics are not reliably safe, and responses that reframe ambiguous requests into safer tasks are substantially less likely to cross the safety boundary. Our results suggest that safe completion should be evaluated as intent-calibrated behavior over controlled task variants, not as a single safety-helpfulness tradeoff over independent prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。