arXiv:2604.12875cs.AI2026-04

195个AI安全评测基准暴露测量标准混乱,缺乏统一语言和维护机制。

AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance

  • 构建195个基准的多层元数据目录,追踪评测设计与指标定义
  • 94个基准属中等复杂度,仅7个被广泛使用,评估高度集中于英文资源
  • 揭示常见指标如准确率背后存在不同评判标准,导致结果不可比

大规模语言模型安全评估的快速发展催生了丰富的评测基准体系,但未形成相应的统一测量体系。本文提出AISafetyBenchExplorer,一个涵盖2018至2026年间发布的195个AI安全评测基准的结构化目录,采用多表架构记录基准级元数据、指标级定义、论文元数据及仓库活动信息。该设计支持对评测存在性及其安全操作化、聚合与评判方式的元分析。分析发现,基准数量激增已远超测量标准化进程:中等复杂度基准占94个,仅7个进入主流层级;165个仅限英文评估,170个为纯评估资源,137个GitHub仓库已停更,96个Hugging Face数据集失效;多数有来源的基准依赖arXiv预印本发布。在指标层面,准确率、F1值、安全评分等常见标签常隐含不同评判主体、聚合规则与威胁模型。我们指出,领域核心问题并非基准稀缺,而是碎片化。研究者虽拥有大量评测资产,却缺乏共享的测量语言、合理的基准选择依据以及持续维护规范。AISafetyBenchExplorer通过可追溯的目录、受控元数据模板与复杂度分类体系,支持更严谨的基准发现、比较与元评估。

原文摘要 · Abstract (English)

The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent measurement ecosystem. We present AISafetyBenchExplorer, a structured catalogue of 195 AI safety benchmarks released between 2018 and 2026, organized through a multi-sheet schema that records benchmark-level metadata, metric-level definitions, benchmark-paper metadata, and repository activity. This design enables meta-analysis not only of what benchmarks exist, but also of how safety is operationalized, aggregated, and judged across the literature. Using the updated catalogue, we identify a central structural problem: benchmark proliferation has outpaced measurement standardization. The current landscape is dominated by medium-complexity benchmarks (94/195), while only 7 benchmarks occupy the Popular tier. The workbook further reports strong concentration around English-only evaluation (165/195), evaluation-only resources (170/195), stale GitHub repositories (137/195), stale Hugging Face datasets (96/195), and heavy reliance on arXiv preprints among benchmarks with known venue metadata. At the metric level, the catalogue shows that familiar labels such as accuracy, F1 score, safety score, and aggregate benchmark scores often conceal materially different judges, aggregation rules, and threat models. We argue that the field's main failure mode is fragmentation rather than scarcity. Researchers now have many benchmark artifacts, but they often lack a shared measurement language, a principled basis for benchmark selection, and durable stewardship norms for post publication maintenance. AISafetyBenchExplorer addresses this gap by providing a traceable benchmark catalogue, a controlled metadata schema, and a complexity taxonomy that together support more rigorous benchmark discovery, comparison, and meta-evaluation.

AI安全评测基准元分析治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。