arXiv:2503.04299cs.AI2025-03被引 7

用现有评测数据帮专家估算大模型实际风险,填补能力与危害间的空白。

Mapping AI Benchmark Data to Quantitative Risk Estimates Through Expert Elicitation

  • 基于Cybench评测数据,邀请专家给出风险发生概率
  • 首次实现从模型能力指标到具体风险概率的量化映射
  • 适合从事AI风险评估、政策制定的研究者参考

大量文献和多位专家指出大型语言模型(LLMs)存在诸多潜在风险,但目前对实际危害的直接测量仍极为有限。现有的AI风险评估主要聚焦于模型能力的测量,而能力仅是风险的指标,并非风险本身。更优的风险建模与量化可弥合这一断层,将大模型的能力与真实世界中的具体损害联系起来。本文在此领域做出初步贡献,展示了如何利用现有AI基准测试促进风险估计的生成。我们报告了一项试点研究的结果:专家们借助Cybench这一AI基准的数据生成了风险概率估计。结果显示该方法具有潜力,同时指出了可改进之处,以进一步提升其在定量AI风险评估中的应用价值。

原文摘要 · Abstract (English)

The literature and multiple experts point to many potential risks from large language models (LLMs), but there are still very few direct measurements of the actual harms posed. AI risk assessment has so far focused on measuring the models' capabilities, but the capabilities of models are only indicators of risk, not measures of risk. Better modeling and quantification of AI risk scenarios can help bridge this disconnect and link the capabilities of LLMs to tangible real-world harm. This paper makes an early contribution to this field by demonstrating how existing AI benchmarks can be used to facilitate the creation of risk estimates. We describe the results of a pilot study in which experts use information from Cybench, an AI benchmark, to generate probability estimates. We show that the methodology seems promising for this purpose, while noting improvements that can be made to further strengthen its application in quantitative AI risk assessment.

AI风险量化评估专家判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。