arXiv:2605.15118cs.CRcs.CL2026-05

建立评估大模型攻击覆盖范围的通用框架,发现现有基准测试仅覆盖不足1/4威胁面。

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

论文配图:Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
图 1 · 摘自论文原文
  • 基于STRIDE构建4×6攻击矩阵,系统梳理932篇论文中的推理时攻击类型
  • 六大数据集仅覆盖矩阵25%单元格,服务中断与模型内部攻击无标准评测
  • 揭示攻击命名混乱(单类攻击最多29种表述)及安全绕过类攻击过度集中

本文提出一个可复用的审计框架,用于评估大模型攻击基准是否全面覆盖威胁面:基于STRIDE构建4×6的目标×技术矩阵,源自包含507个节点的分类体系——其中401个节点由数据填充,106个节点来自威胁建模——从932篇2023至2026年arXiv安全研究中提取的推理时攻击。该矩阵支持基准外验证,聚焦整体覆盖度而非单一基准一致性。应用于六个公开基准发现,三大主流框架(HarmBench、InjecAgent、AgentDojo)占据互不重叠的单元格,累计覆盖不超过25%;而服务中断、模型内部等STRIDE类别在基准中完全缺失,尽管已有相关攻击实现46倍令牌放大和96%成功率,但未被任何基准测试。整个2,521个唯一攻击组揭示命名碎片化严重(单一攻击最多29种表面形式),且大量集中在安全与对齐绕过领域,其结构性特征在小规模下难以察觉。该分类体系、攻击记录与覆盖映射已作为可扩展资源发布;新基准可映射至同一矩阵,帮助社区追踪评估缺口是否缩小。

原文摘要 · Abstract (English)

We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4$\times$6 Target $\times$ Technique matrix grounded in STRIDE, constructed from a 507-leaf taxonomy -- 401 data-populated and 106 threat-model-derived leaves -- of inference-time attacks extracted from 932 arXiv security studies (2023--2026). The matrix enables benchmark-external validation -- auditing collective coverage rather than individual benchmark consistency. Applying it to six public benchmarks reveals that the three primary frameworks (HarmBench, InjecAgent, AgentDojo) occupy non-overlapping cells covering at most 25\% of the matrix, while entire STRIDE threat categories (Service Disruption, Model Internals) lack any standardized evaluation, despite published attacks in these categories achieving 46$\times$ token amplification and 96\% attack success rates through mechanisms which no benchmark tests. The corpus of 2,521 unique attack groups further reveals pervasive naming fragmentation (up to 29 surface forms for a single attack) and heavy concentration in Safety \& Alignment Bypass, structural properties invisible at smaller scale. The taxonomy, attack records, and coverage mappings are released as extensible artifacts; as new benchmarks emerge, they can be mapped onto the same matrix, enabling the community to track whether evaluation gaps are closing.

大模型安全攻击审计威胁建模基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。