arXiv:2412.15433cs.AIcs.CY2024-12被引 3

建立量化模型,预警危险AI能力的出现时机。

Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

  • 构建可追踪危险能力的量化模型,支持政策制定。
  • 模拟显示测试延迟会导致危险评估滞后或偏差增大。
  • 适合关注AI治理与风险预警的研究者和政策制定者。

我们提出一个定量模型,用于追踪危险人工智能能力随时间的变化。该模型旨在帮助政策与研究界可视化危险能力测试如何提前预警接近的AI风险。决策者常依据对AI系统危险性的估计制定政策,并可能设定危险阈值作为政策触发条件。该模型有助于分析此类政策选择。通过模拟,我们揭示了危险能力测试失败的两种表现:对AI危险性估计的偏高,以及阈值监测的显著延迟。其根源在于对AI能力发展动态的不确定性,以及前沿AI实验室间的竞争。有效的AI政策需应对这些失败模式及其驱动因素。即使资源最优配置困难,测试延迟也会损害政策效果。本文提出初步建议,以构建有效的危险能力测试生态系统,并规划相关研究方向。

原文摘要 · Abstract (English)

We present a quantitative model for tracking dangerous AI capabilities over time. Our goal is to help the policy and research community visualise how dangerous capability testing can give us an early warning about approaching AI risks. We first use the model to provide a novel introduction to dangerous capability testing and how this testing can directly inform policy. Decision makers in AI labs and government often set policy that is sensitive to the estimated danger of AI systems, and may wish to set policies that condition on the crossing of a set threshold for danger. The model helps us to reason about these policy choices. We then run simulations to illustrate how we might fail to test for dangerous capabilities. To summarise, failures in dangerous capability testing may manifest in two ways: higher bias in our estimates of AI danger, or larger lags in threshold monitoring. We highlight two drivers of these failure modes: uncertainty around dynamics in AI capabilities and competition between frontier AI labs. Effective AI policy demands that we address these failure modes and their drivers. Even if the optimal targeting of resources is challenging, we show how delays in testing can harm AI policy. We offer preliminary recommendations for building an effective testing ecosystem for dangerous capabilities and advise on a research agenda.

AI治理风险预警测试模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。