提出新基准,测试模型如何应对隐蔽的分步攻击
Benchmarking Misuse Mitigation Against Covert Adversaries
- 构建自动化数据生成流程,模拟分步隐蔽攻击
- 发现弱模型难以拒绝高危任务,而强模型可识别并拒绝
- 验证状态感知防御能有效应对此类复杂攻击
现有语言模型安全评估多针对显式攻击和低风险任务。现实中,攻击者可通过在多个独立查询中请求看似无害的小任务,绕过现有防护机制。单个请求不具威胁性,难以检测,但组合后可实现高危滥用。为识别对此类策略的防御方案,我们开发了状态感知防御基准(BSD),一个自动化评估隐蔽攻击与对应防御的数据生成管道。基于该流程,我们构建了两个新数据集:前沿模型一致拒绝,而较弱的开源模型无法处理。这使得我们能够评估分解攻击的有效性,发现其是典型的滥用助推器,并强调状态感知防御是一种有前景的应对策略。
原文摘要 · Abstract (English)
Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing safeguards by requesting help on small, benign-seeming tasks across many independent queries. Because the individual queries do not appear harmful, the attack is hard to detect. However, when combined, these fragments uplift misuse by helping the attacker complete hard and dangerous tasks. Toward identifying defenses against such strategies, we develop Benchmarks for Stateful Defenses (BSD), a data generation pipeline that automates evaluations of covert attacks and corresponding defenses. Using this pipeline, we curate two new datasets that are consistently refused by frontier models and are too difficult for weaker open-weight models. This enables us to evaluate decomposition attacks, which are found to be effective misuse enablers, and to highlight stateful defenses as a promising countermeasure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。