AI测试需考虑系统可能的策略性行为,否则评估结果不可靠。
AI Testing Should Account for Sophisticated Strategic Behaviour
- 将博弈论引入测试设计,分析AI在评估中的策略性推理
- 实证表明当前评估忽略策略行为会低估风险
- 适合关注AI安全与评估可信度的研究者
本文主张两个观点:第一,为准确反映实际部署表现,AI评估必须考虑系统理解自身处境并进行策略性推理的可能性;第二,博弈论分析可通过形式化和检验基于评估的安全论证中的推理逻辑,优化评估设计。结合现有AI系统的案例、相关研究回顾以及对简化评估场景的形式化战略分析,本文提供了支持这些主张的证据,并提出了若干研究方向。
原文摘要 · Abstract (English)
This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circumstances and reason strategically. Second, game-theoretic analysis can inform evaluation design by formalising and scrutinising the reasoning in evaluation-based safety cases. Drawing on examples from existing AI systems, a review of relevant research, and formal strategic analysis of a stylised evaluation scenario, we present evidence for these claims and motivate several research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。