现有安全评测对小模型效果不佳,模糊判断普遍且影响排名可靠性。
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

- 用统一评分标准测试5个主流安全基准对26个小型模型的评估效果。
- 模糊判断占比高,与提示复杂度和模型结构相关,影响评测可信度。
- 评测结果易受模糊处理方式影响,不适合依赖平均分排序模型。
小型语言模型(SLMs)在资源受限、隐私敏感的场景中日益普及,其安全与偏见问题可能带来安全和社会风险。然而,现有的AI安全/安全/合规评测基准多为大模型设计,未必适用于小模型。本文通过统一评分规则(0=有害,1=安全,0.5=模糊/无关),在26个开源小模型上大规模评估了五个主流基准套件的有效性与鲁棒性。结果显示,模糊判断占主导,且与提示复杂度、模型架构相关,表明‘以大模型为中心的安全基准’不能作为小模型安全评估的独立依据。模糊率随词汇密度、输出困惑度和输出长度上升,随词汇丰富度、自一致性及回复-提示相似度下降。这揭示了能力与安全之间的混淆现象。由于模糊判断普遍,基于均值的排行榜在数学上极为脆弱:即使输出不变,合理处理模糊情况也会显著改变模型排名。
原文摘要 · Abstract (English)
Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-scale assessment of the effectiveness and robustness of these automated pipelines by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which assigns a score of 0, 1, or 0.5 to harmful, safe, or ambiguous/irrelevant responses, respectively. Across the benchmarks, ambiguous judgments dominate and correlate with prompt complexity and model architecture, indicating that {\em LLM-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment}. In general, the ambiguity rate increases with lexical density, output perplexity, and output length and decreases with lexical sophistication, self-coherence, and reply-prompt similarity. This reveals a capability-safety confound that mixes model capability with apparent safety. Since ambiguity is prevalent, aggregate mean-score leaderboards are mathematically brittle: model rankings change significantly under reasonable ambiguity treatments, even when the underlying outputs remain unchanged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。