arXiv:2607.13039cs.CYcs.AI2026-07被引 1

评估生物助手双重用途风险防护机制的有效性与用户影响。

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

  • 区分拒绝行为来源,按干预预算优化防护阈值。
  • 多组实验显示仅部分配置满足合法用户保护要求。
  • 揭示验证器界面设计对风险控制的关键影响。

现有拒绝率无法识别干预环节或衡量对合法用户的负担。本文在操作和回答两个层面评估双用途生物助手的安全防护机制。框架重构访问路径,分离提供方拒绝与下游动作,并在干预预算下选择阈值。冷冻生成研究显示,Claude Opus 4.5 满足其标准,但两种通过配置均受同一上游提供方影响;而 Gemini 2.5 Flash 无一通过。固定 Opus 策略在 104 个未使用过的标注对上保持正向选择性,但均不满足 20% 匹配良性样本约束。在回答层面,7,200 条独立判断的联合评分验证器均未达标。一项新开展的 8,640 判断因子实验发现,标准隔离与序数表示分别带来提升,且二者存在正向交互;但在严格不可修复框架下,显式定位会降低整体准确率。证据支持前瞻性操作级选择性,揭示验证器界面效应,但未验证校准选择性接入、内容移除或生物风险降低。

原文摘要 · Abstract (English)

A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the action and answer levels. The framework reconstructs the access path, separates provider refusals from downstream actions, and selects thresholds under an intervention budget. A frozen fresh-generation study satisfies its criterion on Claude Opus 4.5, but both passing configurations share one upstream provider effect; none passes on Gemini 2.5 Flash. The fixed Opus policies retain positive selectivity on 104 previously unused released-label pairs, but both fail a 20\% matched-benign constraint. At the answer level, no joint-scoring verifier qualifies on a response-disjoint 7,200-judgment holdout. A fresh 8,640-judgment factorial experiment finds separate gains from criterion isolation and ordinal representation, with a positive interaction between them; requiring explicit localization lowers aggregate accuracy under a strict no-repair schema. The evidence supports prospective action-level selectivity and identifies verifier interface effects, but not calibrated selective access, verified content removal, or biological-risk reduction.

生物安全风险评估模型防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。