不同监管策略会影响AI风险治理,优先级与多样性平衡更有效
Supervision policies can shape long-term risk management in general-purpose AI models
- 设计仿真框架,测试四种风险监管策略的实效
- 专家识别的高影响风险在优先与多元策略下更易被控制
- 忽视社区报告可能引发风险感知失真,适合政策制定者参考
通用人工智能(GPAI)模型的快速普及给监管机构带来前所未有的挑战。我们假设监管方将面临远超其能力的风险报告生态。为此,构建了一个基于多元风险报告系统特征的仿真框架,涵盖社区平台、众包和专家评估。评估四种监管策略:非优先(先到先处理)、随机选择、优先级(优先处理高风险)和多样性优先(兼顾高风险与类型覆盖)。结果表明,优先级和多样性策略能更有效缓解专家识别的高影响风险,但可能忽略来自广泛社区的系统性问题,导致反馈循环,加剧某些报告类型而抑制其他类型,扭曲整体风险认知。通过包含超过一百万次ChatGPT交互的真实数据集验证,其中逾15万次对话被标记为高风险。该研究揭示了AI风险监管中的复杂权衡,说明监管策略选择将深刻塑造未来社会中各类GPAI模型的风险格局。
原文摘要 · Abstract (English)
The rapid proliferation and deployment of General-Purpose AI (GPAI) models, including large language models (LLMs), present unprecedented challenges for AI supervisory entities. We hypothesize that these entities will need to navigate an emergent ecosystem of risk and incident reporting, likely to exceed their supervision capacity. To investigate this, we develop a simulation framework parameterized by features extracted from the diverse landscape of risk, incident, or hazard reporting ecosystems, including community-driven platforms, crowdsourcing initiatives, and expert assessments. We evaluate four supervision policies: non-prioritized (first-come, first-served), random selection, priority-based (addressing the highest-priority risks first), and diversity-prioritized (balancing high-priority risks with comprehensive coverage across risk types). Our results indicate that while priority-based and diversity-prioritized policies are more effective at mitigating high-impact risks, particularly those identified by experts, they may inadvertently neglect systemic issues reported by the broader community. This oversight can create feedback loops that amplify certain types of reporting while discouraging others, leading to a skewed perception of the overall risk landscape. We validate our simulation results with several real-world datasets, including one with over a million ChatGPT interactions, of which more than 150,000 conversations were identified as risky. This validation underscores the complex trade-offs inherent in AI risk supervision and highlights how the choice of risk management policies can shape the future landscape of AI risks across diverse GPAI models used in society.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。