测试AI在生物安全任务中识别风险与拒绝违规请求的能力
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
- 构建配对任务集:61项常规科研任务+46项伪装成科研的生物安全威胁
- 部分模型对合法任务拒绝对比隐蔽威胁更高,最高达74%拒绝率
- 模型推理空间越大,越能识别真实威胁,适合开发者评估AI安全
随着AI代理被引入生命科学流程,其加速发现的能力也可能被滥用。我们提出BioSecBench-Refusal,一个用于生物科研任务中风险识别与拒绝行为的基准测试。该基准包含61个常规任务(源自已发表文献的合法分析)和46个红队任务(看似真实研究但隐藏生物安全风险的虚构场景)。在16种模型配置下,常规任务的拒绝率在7%至74%之间,红队任务为1%至62%,许多配置对合法任务的拒绝率与隐蔽威胁相当甚至更高。拒绝对主要由代理前的API过滤器触发,但给予模型足够推理空间后,其识别真实威胁的能力显著提升。我们发布BioSecBench-Refusal,供模型开发者在智能生物技术研发中平衡能力与安全。
原文摘要 · Abstract (English)
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7 percent to 74 percent on Routine tasks and 1 percent to 62 percent on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech research and development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。