提出国际AI事故升级标准,解决何时需跨国协作的难题。
Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds
- 构建八项评估标准的分级决策流程,明确升级触发条件。
- 测试发现三类设计缺陷导致事故漏报,如依赖已确认伤害。
- 强调定义与数据基础对标准有效性的影响,适合政策制定者参考。
AI事故报告要求正逐步纳入法规与政策,但尚无操作性标准来判断何种事件需从国家层面升级至国际协调。本文提出一个升级框架,旨在为各司法管辖区提供共同参考,实现协同升级的同时保留各主体在法律和政策框架内的响应灵活性。通过分析SB 53、欧盟《人工智能法案》、GPAI行为准则及其他行业事故框架,提炼出八项评估标准,并转化为带分步审查点与阈值校验的流程图。针对十起已记录的AI事故及结构化变体进行测试,识别出标准在实践中可能漏检或误判的情形。研究发现三类系统性漏检模式:(a)要求确认实际损害,导致模型权重泄露等风险仅在严重不可逆后果发生后才被识别;(b)按单个事件评估,使累积性系统性危害被低估;(c)阈值与法律工具绑定而非可量化检验,造成紧急情境下难以应用。此外,升级规则仅为整体框架的一部分,底层定义与可用数据也构成关键依赖,可能引发漏报。
原文摘要 · Abstract (English)
AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI incident warrants escalation beyond national handling to international coordination. This paper proposes an escalation framework to address this gap, intended as a common reference point across jurisdictions that enables aligned escalation while preserving flexibility in how actors respond within their own legal and policy contexts. We review SB 53, the EU AI Act, the GPAI Code of Practice, and incident frameworks from other industries to derive eight criteria for assessing whether an incident warrants escalation, translated into a sequential flowchart with gated decision points and threshold checks. For each criterion, we map how it interplays with these regulatory frameworks, identifying where their design choices support or undermine effective detection. We test the framework against ten documented AI incidents and structured variants to identify where criteria under-detect or misclassify incidents in practice. We find three design patterns that may lead to systematic under-detection in regimes where model developers are responsible for escalation: a. where escalation requires confirmed harm, events such as model weight exfiltration risk detection only after severe, irreversible harm has propagated; b. where incidents are assessed individually, systemic harms emerging from accumulation risk being under-detected; and c. where thresholds align with legal instruments rather than quantitatively testable terms, criteria risk being impractical to apply under time pressure. We also find that escalation rules are only one component of a broader framework: the underlying definitions against which thresholds are set, and the data available to the responsible actor, create interdependencies that can themselves drive under-detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。