现有行为验证无法证明AI安全,因它看不见深层机制。
Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

- 用行为测试和红队评估验证安全,但看不到模型内部表示
- 实测中80%的安全声明缺乏可验证证据支持
- 适合政策制定者与合规团队参考
本文指出,尽管精心设计,行为保证仍被要求承担其无法验证的安全承诺。2019至2026年间出台的AI治理框架要求提供可审查证据,证明模型无隐藏目标、抗失控前兆、能力受限等属性;然而当前主要依赖行为评估与红队测试的方法,仅能观察输出结果,无法验证潜在表征或长期自主行为。我们提出“审计缺口”概念,即期望验证范围与实际可访问信息之间的结构性不匹配,并引入“脆弱保证”描述证据结构无法支撑安全主张的情形。通过对21项工具的分析,发现地缘政治与产业压力正系统性激励表面行为代理,而非深层结构验证。最后建议技术转向:在法律文本中限制行为证据权重,并通过线性探针、激活修补、训练前后对比等机制证据,扩展自愿部署前的可访问性。
原文摘要 · Abstract (English)
This position paper argues that behavioural assurance, even when carefully designed, is being asked to carry safety claims it cannot verify. AI governance frameworks enacted between 2019 and early 2026 require reviewable evidence of properties such as the absence of hidden objectives, resistance to loss-of-control precursors, and bounded catastrophic capability; current assurance methodologies (primarily behavioural evaluations and red-teaming) are epistemically limited to observable model outputs and cannot verify the latent representations or long-horizon agentic behaviours these frameworks presume to regulate. We formalize this structural mismatch as the audit gap, the divergence between required and achievable verification access, and introduce the concept of fragile assurance to describe cases where the evidential structure does not support the asserted safety claim. Through an analysis of a 21-instrument inventory, we identify an incentive gradient where geopolitical and industrial pressures systematically reward surface-level behavioral proxies over deep structural verification. Finally, we propose a technical pivot: bounding the weight of behavioral evidence in legal text and extending voluntary pre-deployment access with mechanistic-evidence classes, specifically linear probes, activation patching, and before/after-training comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。