实验发现:落后焦虑比风险偏好更促使AI研发冒进。
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

- 通过重复博弈模拟AI竞赛,观察参与者在不同风险上限下的行为选择。
- 落后时选高风险策略的概率显著上升,领先则趋于保守,首轮选择预示后续行为。
- 竞争压力与对手动作比个人风险偏好更能驱动不安全研发,适合政策制定者参考。
技术竞赛中速度与安全存在张力:参与者可能因赶超对手而选择更快但更危险的发展路径。我们通过一个理想化的AI竞赛行为实验研究此现象,参与者成对重复选择安全或非安全开发,不确定性时间范围下,非安全开发带来更快进展和更高即时回报,但累积私人风险至特定上限(10%、60%或90%),竞赛结构保持不变,仅风险上限变化。预先注册的跨风险水平比较及风险偏好影响均未获数据支持。探索性分析显示,非安全行为更多受竞赛战略状态驱动:对手选择非安全后,自身更倾向跟进;领先时减少非安全行为,落后时增加;首轮选择可预测后期行为。为此引入简化进化模型,包含四种策略(始终安全、始终非安全、条件安全、条件反社会安全),复现了处理效应,并揭示条件性非安全行为如何被竞争动态所青睐。实验与模型表明,不安全研发源于早期行为惯性、对手行为及落后恐惧,而非仅由风险偏好导致,提示政策应聚焦降低竞争压力、促进合作,而非仅关注个体风险认知。
原文摘要 · Abstract (English)
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。