用心理测量学方法提升大模型安全评估的可信度与效率
Item Response Theory for AI Safety

- 引入项目反应理论,从多基准中提取模型安全性的可解释维度
- 仅需约10个精选题目即可高精度还原完整评估结果,成本降低97%-99%
- 能检测模型伪装行为和接口后模型变更,适合安全审计与前沿实验室使用
语言模型在安全性上的差异通过安全基准进行衡量,但聚合评分难以信赖且难以解释,因为基准间重复性高、高度相关,且模型可能在检测到评估时采取策略性回避。为解决这些问题,我们引入项目反应理论(IRT),一种用于从带有推断心理测量属性的题目表现中测量潜在特质的统计工具。我们在192个语言模型上对8个安全基准拟合了IRT模型,这是迄今为止对大模型安全评估最全面的心理测量分析。研究发现:三个可解释因素——拒绝严格性、真实性与情境危害性,解释了模型间大部分差异;经心理测量筛选的题目,其恢复完整基准分数的误差低于随机选取相同数量题目的情况,约10个自适应选择的题目即可实现多个基准的高精度评估,使评估成本降低97%-99%;此外,IRT可用于单个模型的审计,能够识别出简单的伪装行为及接口后模型变化。总体而言,IRT是一个现成的工具包,可用于解读、简化和审计安全基准,我们建议前沿实验室和评估者采用。
原文摘要 · Abstract (English)
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。