检测大模型拒答是否真可靠,发现多数模型拒答经不起提示微调就失效。
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

- 用稀疏自编码器分析模型内部激活,评估拒答是否真实可靠。
- Gemma 2 和 Gemma 4 模型在不同提示下拒答率从 65% 跌至 0%,80 字限制下全失效。
- 部分模型将无害生物物质误判为危险,拒答更依赖法律和文化因素而非安全风险。
语言模型生物安全评估通常关注其是否生成有害内容。本文提出一个互补问题:当模型拒绝时,这种拒绝是结构性稳健的,还是在提示框架、格式或输出长度稍作修改后就消失?在五种架构中,没有模型能清晰区分良性与危险内容。Gemma 2 2B-IT 在 75 个提示中从未真正拒绝,对每个接近危险的查询都含糊其辞;Gemma 4 E2B-IT 在聊天模板格式下拒绝了 65/75 个提示,无格式时拒绝率为 0/75。两者在 80-token 长度限制下拒答率均降为 0%。Qwen 2.5 1.5B 与 Phi-3-mini 过度拒绝,将 83%-87% 的良性生物学内容标记为危险。只有 Llama 3.2 1B 显现出有意义的层级差异(61 分差)。为探究过拒原因,测试了若干列管一级但生物学无毒化合物(如具 FDA 突破性疗法地位的裸盖菇素培养),部分模型对这些内容的拒绝率甚至超过真正危险生物,表明拒答更依赖合法性与文化敏感性而非 CBRN 危险性。为此引入分歧度量 D,比较模型表面响应标签与其内部稀疏自编码器(SAE)特征激活。在 Gemma 2 2B-IT(Gemma Scope 1)与 Gemma 4 E2B-IT(作者训练生物领域 SAE)上计算完整 D 值。释放两个微调后的 Gemma 2 域 SAE。在 Gemma 4 上,合规与拒绝响应间激活差距达 0.647 分,无重叠(n=75),尽管仍属初步结果,样本有限、校准局限且仅覆盖 Gemma 家族。该研究于一周内使用消费级硬件(GTX 1650 Ti Max-Q 及 Colab T4 训练 SAE)完成,初步表明激活层面审计可揭示行为评估无法察觉的失效模式,且各架构表现差异显著。
原文摘要 · Abstract (English)
Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds (notably psilocybin cultivation, with FDA Breakthrough Therapy status). Some models refused these at rates exceeding genuinely hazardous biology, suggesting refusal tracks legality and cultural salience over CBRN hazard. To measure the internal side, we introduce a divergence score D comparing a model's surface response label to its internal sparse autoencoder (SAE) feature activations. Full D was computed on Gemma 2 2B-IT (Gemma Scope 1) and Gemma 4 E2B-IT (author-trained bio SAE). Two fine-tuned Gemma 2 domain SAEs were released. On Gemma 4, comply and refuse responses separated by a 0.647-point gap with zero overlap (n=75), though this is preliminary, with a narrow catalog, within-sample calibration, and Gemma-family-only SAE coverage. Built over one hackathon weekend on consumer hardware (GTX 1650 Ti Max-Q, plus Colab T4 for SAE training), this preliminary evidence suggests activation-level auditing may surface failure modes invisible to behavioral evaluation, with substantial variation across architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。