arXiv:2503.20848cs.GTcs.AI2025-03被引 8

弱监管反而降低AI安全,强而精准的监管能提升所有参与者利益

The Backfiring Effect of Weak AI Safety Regulation

  • 构建三方博弈模型,分析监管、通用AI开发者与应用专家的互动
  • 对领域专家施加弱监管会降低整体安全水平,效果适得其反
  • 对通用与应用双端施加强监管可提升安全与性能,实现共赢

近期政策提案试图提升通用型AI的安全性,但对不同监管方式的有效性缺乏理解。本文提出一个战略模型,探讨安全监管、通用型AI开发者及领域专家(将技术适配于具体应用场景者)之间的互动。分析聚焦于不同监管措施针对AI开发链不同环节的影响。假设AI技术具有安全性和性能两个关键属性:监管方首先设定最低安全标准,适用于一方或双方,违规将受严格处罚;通用开发者随后投资研发,确立初始安全与性能水平;领域专家再针对特定场景优化,更新安全与性能,并推向市场;收益按收入共享参数在双方间分配。分析揭示两个核心发现:第一,主要针对领域专家的弱监管可能产生反效果。尽管看似应监管使用场景,但分析表明,仅对领域专家施加弱监管会意外降低整体安全性,且该效应在多种情境下持续存在。第二,与前述相反,适当且有力的监管可使所有受监管方获益。当监管方对通用型开发者和领域专家同时施加合理安全标准时,监管成为承诺机制,推动安全与性能双重提升,超越无监管或仅监管一方的情形。

原文摘要 · Abstract (English)

Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches to AI safety. We present a strategic model that explores the interactions between safety regulation, the general-purpose AI creators, and domain specialists--those who adapt the technology for specific applications. Our analysis examines how different regulatory measures, targeting different parts of the AI development chain, affect the outcome of this game. In particular, we assume AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety standard that applies to one or both players, with strict penalties for non-compliance. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their specific use cases, updating the safety and performance levels and taking the product to market. The resulting revenue is then distributed between the specialist and generalist through a revenue-sharing parameter. Our analysis reveals two key insights: First, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can unintentionally reduce safety. This effect persists across a wide range of settings. Second, in sharp contrast to the previous finding, we observe that stronger, well-placed regulation can in fact mutually benefit all players subjected to it. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation functions as a commitment device, leading to safety and performance gains, surpassing what is achieved under no regulation or regulating one player alone.

AI治理监管策略博弈模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。