AI外交决策风险高,现有评估方法难应对,亟需新评价框架。
The Foreign Policy AI Evaluation Gap

- 将外交任务分解为可评估的子任务,结合人工重组提升可测性。
- 外交领域存在不可观测、目标多元等特性,传统评估方法失效。
- 适合关注AI治理、国际关系与技术安全的研究者参考。
我们主张,用于执行外交任务(广义上的国家权术)的AI系统应成为技术性AI治理研究的优先案例。外交涉及政治主体制定并实施对外目标,属于高后果部署领域,具有极端下行风险,且其结构性特征使得标准评估实践难以应对:包括部分可观测性、无界行动空间、争议性真实标签及多维目标。本文提出三项贡献:(i) 指出外交领域兼具灾难性尾部风险与技术评估复杂性的结构条件;(ii) 通过生态审查揭示当前研究对评估功能的关注远超访问、验证、安全与操作化;(iii) 提出需求侧评估框架,将外交工作流拆解为边界明确、可评估的子任务,并允许人工重组。随着AI已在战争与和平事务中实际部署,而技术型AI治理社区缺乏公共评估基础设施,该研究议程具有紧迫性。
原文摘要 · Abstract (English)
We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for technical AI governance research. In enacting foreign policy, we refer to the formulation and implementation of external objectives by political actors. Statecraft is a high-consequence deployment domain, with extreme downside risks and structural properties that standard evaluation practices handle poorly. These features include partial observability, unbounded action spaces, contested ground truth, and multidimensional objectives. This paper advocates for a literature-grounded research agenda. Our contribution is threefold: (i) a claim about the structural conditions of foreign policy that combine catastrophic tail risk with technical evaluation complexities, (ii) an ECOSYSTEM review that highlights the asymmetric focus on ASSESSMENT features over ACCESS, VERIFICATION, SECURITY, and OPERATIONALIZATION, and (iii) a demand-side evaluation framework that decomposes foreign-policy workflows into bounded, evaluable sub-tasks with human recombination. As AI systems are already being deployed in the conduct of war and peace, amid limited public evaluation infrastructure from the technical AI governance community, this agenda is an urgent priority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。