构建评估大模型危险能力的统一框架,揭示安全与性能并非同步提升。
FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

- 分知识、防御、危害三维度,用标准化流程评估模型危险性
- 12款主流模型测试显示:新模型知识更强但防御改进有限
- 适合关注AI安全治理、模型风险评估的研究者和从业者
零散的安全评估阻碍了危险人工智能能力的治理。我们提出一个模块化框架,通过知识(K)、防御(D)、危害(H)三个正交管道,在统一协议下评估模型,并聚合为标准化的危险能力画像ϕ。可插拔模块提供场景种子、知识库、危险问题和评判标准,核心评估引擎保持跨领域一致;CB评估通过网络模拟验证了协议迁移性。使用化学生物(CB)模块评估来自四个家族的12款商用大模型。第一,横向比较揭示各模型与家族间危险能力差异显著:知识相近的模型在拒绝韧性上表现不同,强防御模型在合规时仍可能生成更多有害内容;家族层面呈现Claude、DeepSeek、GPT模型的清晰分异。第二,时间演化分析显示:随着模型发布时间推移,知识(K)持续增强,防御(D)仅部分改善,危险能力未单调下降,说明规模扩展与对齐进展并未均匀带来安全性提升。可靠性通过跨评卷人一致性(自助法ρ > 0.79,5名评委中4名)与管道正交性(K-D-H间相关系数ρ ∈ [0.32, 0.52])得到验证。
原文摘要 · Abstract (English)
Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $\phi$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking $K$, $D$, and $H$ against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap $\rho > 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$--$D$--$H$ inter-correlations $\rho \in [0.32, 0.52]$).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。