arXiv:2606.20626cs.CYcs.AI2026-06

用心理测量学方法让安全评估效率提升99%,少跑大量测试仍准。

Efficient Safety Benchmarking via Item Response Theory

论文配图:Efficient Safety Benchmarking via Item Response Theory
图 1 · 摘自论文原文
  • 引入项目反应理论,让模型安全能力可量化对比。
  • 自适应选题只需原成本20%以下,就能逼近完整评测结果。
  • 提取通用高信息量题目,适合批量测试与长期评估。

语言模型的安全评测通常采用静态范式,对所有模型一视同仁地使用全部题目,这种假设在对抗性、异质性强的安全题目中尤为不成立。当前全流程评测需约10⁵次响应,其中多数提供极少排序信号。我们分析了多个主流安全评测集,提出三项改进:首先,项目反应理论(IRT)能揭示安全评测的可解释结构,其能力估计可区分在原始安全得分上均达天花板的模型;其次,基于响应动态选择高信息量题目的自适应评测,在达到与全量评测相关系数ρ>90%的基准下,成本降低至少80%,在AIR-Bench 2024上最高可降99.9%;第三,提出一种实用方案,提取固定且高信息量的题目子集,实现跨模型复用,同样在AIR-Bench 2024上节省高达99.8%成本。三者结合表明,心理测量方法可显著降低整个安全评测流程的开销。

原文摘要 · Abstract (English)

Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full to modern benchmark suites, the current evaluation procedures would require on the order of $10^5$ responses, most of which provide little ranking signal. We analyze a suite of widely used safety benchmarks and make three contributions toward more efficient safety evaluation. First, we show that Item Response Theory (IRT) recovers interpretable structure on safety benchmarks, with ability estimates resolving differences among models that cluster at the ceiling of raw safety metrics. Second, we show that adaptive item selection, which dynamically chooses informative items for each model based on its responses, approximates full-benchmark rankings while reducing evaluation cost by at least 80% on benchmarks where Spearman's $ρ>$90% with full-benchmark is attainable, and by up to 99.9% on AIR-Bench 2024. Third, we introduce a practical procedure for extracting a fixed, informative subset of items reusable across models, providing an alternative to adaptive selection with savings of up to 99.8% on AIR-Bench 2024. Together, these results establish that psychometric methods enable benchmark-aware reductions in evaluation costs across the safety evaluation pipeline.

安全评测项目反应理论效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。