提出可重构性度量,让大模型安全评估结果更可信
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric
- 定义八类决策属性的可重构性指标,量化证据能否还原决策
- 实测12个场景下可重构性仅0.458-0.833,多数评估无法复现
- 生成证据充分性卡片,适合评估者与审计人员验证安全性声明
当前许多智能体安全评估结果尚不足以作为支撑性证据:相同的任务成功或攻击成功率可能基于完全不同证据体系。缺乏可跨厂商、可运行的工具来衡量评估有效性——即所采集证据是否足以重构决策依据。本文提出一种基于八类决策属性的可重构性度量,并设计跨评估框架适配器,生成每项决策的证据充分性卡片,用于运行时监控覆盖检查。提出反事实重播干预协议,实现重播前提探测,并定义了声明-证据过充差距。在公开及捆绑数据轨迹上,无需新增模型运行,12个领域内证据充分性跨度为0.458–0.833,四个输入具有相同表面表现;所有评分轨迹均未满足重播前提条件。在合成发布门控对中,原始版本(0.542)被阻断,而工具化版本(0.667)通过。安全评估结论应附带其可重构性向量,复现包可再生所有报告数值。
原文摘要 · Abstract (English)
Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores reconstructability as an evaluation-validity metric: whether captured evidence can reconstruct the decision a claim depends on. This paper introduces a property-level reconstructability metric over eight decision-property classes and a cross-harness adapter emitting per-decision Evidence Sufficiency Cards backing a per-run monitor-coverage release check. It specifies a counterfactual-replay intervention protocol, implements its replayability-precondition probe, and defines a claim-evidence overclaim gap. On public and bundled traces, without new model runs, twelve-field sufficiency spans 0.458-0.833 across four inputs sharing a surface reading; replay preconditions are unmet in every scored trace. In a synthetic release-gate pair, the sufficiency gate blocks the raw variant (0.542) and passes the instrumented (0.667). Safety-evaluation claims should travel with their reconstructability vector; a reproducibility package regenerates every reported number.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。