arXiv:2505.17636cs.LGcs.AI2025-05

分析六个主要危害类别在安全评测集中的分布差异,揭示评测标准的隐蔽偏移。

Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis

  • 用UMAP与kmeans识别出五大数据集的语义聚类模式
  • 发现不同评测集对隐私、自伤等六类危害覆盖不均,如GretelAI侧重隐私
  • 提出量化框架,帮助研发者补全评测盲区

我们对五个近期发布的开源模型安全评测集进行了评估,通过UMAP降维和kmeans聚类(轮廓系数0.470)发现了明显的语义簇。识别出六类主要危害,其在不同评测集中的分布存在显著差异:例如GretelAI集中于隐私问题,WildGuardMix则侧重自伤场景。提示长度分布的差异表明数据收集过程可能存在混杂因素,影响对危害的理解。本研究量化了不同评测集间的语义正交性,尽管主题相似,仍能揭示覆盖盲区。所提出的量化框架可支持更精准地开发全面应对未来人工智能危害演化的数据集。

原文摘要 · Abstract (English)

Various AI safety datasets have been developed to measure LLMs against evolving interpretations of harm. Our evaluation of five recently published open-source safety benchmarks reveals distinct semantic clusters using UMAP dimensionality reduction and kmeans clustering (silhouette score: 0.470). We identify six primary harm categories with varying benchmark representation. GretelAI, for example, focuses heavily on privacy concerns, while WildGuardMix emphasizes self-harm scenarios. Significant differences in prompt length distribution suggests confounds to data collection and interpretations of harm as well as offer possible context. Our analysis quantifies benchmark orthogonality among AI benchmarks, allowing for transparency in coverage gaps despite topical similarities. Our quantitative framework for analyzing semantic orthogonality across safety benchmarks enables more targeted development of datasets that comprehensively address the evolving landscape of harms in AI use, however that is defined in the future.

模型安全评测集分析语义正交性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。