arXiv:2605.21545cs.SEcs.AI2026-05被引 1

对比大模型在生物研究中的拒答行为,发现现有评估方法会误判模型安全水平。

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

  • 设计匹配三元组数据集,固定任务框架仅变生物风险等级,避免领域干扰。
  • 19个前沿模型拒答率从0.1%到94.6%,部分模型连应拒提示也不拒绝。
  • 拒答行为受接口路径影响远大于模型本身,体现模板化响应而非深度推理。

前沿大模型正被广泛用于生物研究工作流的调度,但缺乏统一的拒绝行为评估基准。本文提出RefusalBench,一个包含141个提示、47个组合的匹配三元组基准,保持任务框架一致,仅改变生物风险等级(无害、边缘、双用途),实现风险层级条件下的稳健比较。15个应拒提示的正控模块建立各模型校准基线;其中三个模型甚至未拒绝这些提示。在2026年5月快照中的19个前沿模型中,相同提示上的严格拒答率跨度为0.1%至94.6%。司法管辖区无法预测拒答行为(Mann-Whitney U, p = 0.393;欧盟n=1,美国呈双峰分布);而提供方身份具有显著影响,Anthropic API栈可预测拒答(OR = 21.03,95% CI: 14.58–30.34,提示聚类;5.70–77.55,模型聚类GEE)。该效应更倾向接口路径层级而非模型权重层级:Anthropic 99.8%的严格拒答使用相同的政策判定代码,表明其基于少数标准模板,而非逐例推理。严格拒答率会错误排名安全校准水平:Grok 4.20在风险区分度上表现最佳(Youden's J = 0.787),但整体拒答率仅排第七;Claude Opus 4.7的区分度下降65%且双用途检测未提升。18个前沿模型中有9个在双用途层级表现出‘部分合规’模式,而二元拒答指标无法捕捉此现象。

原文摘要 · Abstract (English)

Frontier large language models are increasingly deployed as orchestration backbones for biological research workflows, yet no shared evidence base exists for comparing their refusal behaviour on legitimate research prompts. RefusalBench, introduced here, is a matched-triple benchmark of 141 prompts in 47 bundles that holds task framing constant while varying only biological risk tier (benign, borderline, dual-use), enabling tier-conditioned comparisons robust to subdomain confounding. A 15-prompt should-refuse positive-control module establishes per-model calibration floors; three models fail to refuse even these prompts. Across 19 frontier models in the May 2026 snapshot, strict refusal rates span 0.1% to 94.6% on identical prompts. Jurisdiction does not predict refusal in this snapshot (Mann-Whitney U, p = 0.393; EU n = 1, US bimodal); provider identity does, with Anthropic's API stack predicting refusal at OR = 21.03 (95% CI: 14.58-30.34 prompt-clustered; 5.70-77.55 under model-clustered GEE). This effect is best read as access-path-level rather than model-weight-level: 99.8% of Anthropic's strict refusals carry the same safety_policy adjudicated reason code, consistent with a small set of canonical refusal templates rather than case-by-case model reasoning. Strict refusal rate misranks safety calibration: Grok 4.20 achieves the highest tier discrimination (Youden's J = 0.787) while ranking only seventh by overall refusal rate, and Claude Opus 4.7's J dropped 65% from prior versions with no improvement in dual-use detection. Nine of 18 frontier models exhibit a hedge-but-help partial-compliance pattern at dual-use tier that binary refusal metrics cannot detect.

大模型安全生物研究拒答行为评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。