量化AI研发自动化程度及其对能力与安全的影响
Measuring AI R&D Automation
- 设计多维度指标追踪AI研发自动化水平
- 发现自动化可能使能力进步快于安全进展
- 适合政策制定者与研究机构参考使用
AI研发自动化(AIRDA)可能带来深远影响,但其实际程度及最终后果仍不明确。现有数据(主要为能力基准测试)难以反映真实世界中的自动化情况,也未能捕捉其广泛影响,例如自动化是否使能力提升速度超过安全进展,或监管能力能否跟上研发加速。为此,本文提出一系列指标,用于衡量AIRDA的程度及其对AI进步和监管的效应。这些指标涵盖资本占比、研究人员时间分配、AI越狱事件等维度,有助于决策者理解自动化后果,采取适当安全措施,并保持对AI发展速度的感知。建议企业与第三方组织(如非营利研究机构)开始追踪这些指标,政府应支持此类努力。
原文摘要 · Abstract (English)
The automation of AI R&D (AIRDA) could have significant implications, but its extent and ultimate effects remain uncertain. We need empirical data to resolve these uncertainties, but existing data (primarily capability benchmarks) may not reflect real-world automation or capture its broader consequences, such as whether AIRDA accelerates capabilities more than safety progress or whether our ability to oversee AI R&D can keep pace with its acceleration. To address these gaps, this work proposes metrics to track the extent of AIRDA and its effects on AI progress and oversight. The metrics span dimensions such as capital share of AI R&D spending, researcher time allocation, and AI subversion incidents, and could help decision makers understand the potential consequences of AIRDA, implement appropriate safety measures, and maintain awareness of the pace of AI development. We recommend that companies and third parties (e.g. non-profit research organisations) start to track these metrics, and that governments support these efforts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。