用7714起真实事件数据检验了大模型安全十大风险排名的可靠性。
Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus

- 构建7714起真实事件库,用贝叶斯模型修正分类误差,生成数据驱动排名。
- 专家排名与数据排名相关性弱(κ≈0.20),但专家共识仍具稳健性。
- 验证了当前分类器无法超越基础准确率86.3%,适合关注大模型安全评估者。
OWASP大模型十大风险依赖安全专家共识排序。本文提出更具体问题:该共识是否经得起真实事件数据检验?研究收集了来自CVE、GHSA、OSV和AIAAIC的7714条快照及6639条标注事件,基于20类分类体系,采用贝叶斯测量误差模型校正每类别的计数偏差。2026候选清单以0.75权重融合专家投票与0.25权重数据信号,使数据修正共识但不颠覆它。两排名间一致性较弱(Cohen's κ≈0.20,90%置信区间跨越零点)。专家排名仍具稳健性:预注册的四个前沿分类器比拼无胜者,均未突破基线平衡准确率86.3%。真实标签验证显示基线排序与独立真值高度一致(Spearman ρ=0.918)。本研究由工作组成员开展,非官方发布,不取代正式列表或流程。
原文摘要 · Abstract (English)
The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower question: checked against the record of real incidents, does that expert ranking agree with the data? We assembled a large-scale corpus of LLM-security incidents (7,714 snapshotted and 6,639 labeled against the 20-entry taxonomy) drawn from CVE, GHSA, OSV, and AIAAIC, and derived an incident-based ranking with a Bayesian measurement-error model that corrects each category's count for classifier precision and recall. The 2026 candidate list blends the two signals at fixed weights, 0.75 on the expert vote and 0.25 on the data, so the corpus corrects the consensus without overturning it. The agreement between the two rankings is weak: Cohen's $κ\approx 0.20$, with a 90% interval that crosses zero. The expert ranking is nonetheless robust. A pre-registered bake-off of four frontier classifiers returns no winner. None beats the incidence floor's balanced accuracy of 0.863. A ground-truth check leaves the floor's ordering (Spearman $ρ= 0.918$ against held-out truth) in place. This is an exploratory analysis by two working-group members, not the official OWASP release, and it does not supersede the official list or process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。