AI安全需具备抗脆弱性,让系统随挑战越变越强。
Position: AI Safety Must Embrace an Antifragile Perspective
- 用动态挑战机制替代静态测试,主动暴露模型弱点
- 通过不确定性反哺训练,提升对罕见事件的应对能力
- 适合长期部署的AI系统研发者与安全评估团队
本文主张现代AI研究必须采用抗脆弱性安全视角——即系统处理稀有或分布外(OOD)事件的能力应随时间增强。传统静态基准和一次性鲁棒性测试忽略了环境演变的事实,若模型长期未受挑战,可能产生适应性退化(如奖励黑客、过度优化或能力萎缩)。我们提出,与其急于消除当前不确定性,不如利用这些不确定性来更好应对未来更复杂、更不可预测的挑战。本文首先指出静态测试在场景多样性、奖励黑客和过度对齐方面的关键局限,再探讨抗脆弱方案对罕见事件管理的潜力。核心主张是重新校准衡量、基准测试和持续改进AI安全的方法,以提供伦理与实践指南,推动构建具有抗脆弱性的AI安全共同体。
原文摘要 · Abstract (English)
This position paper contends that modern AI research must adopt an antifragile perspective on safety -- one in which the system's capacity to guarantee long-term AI safety such as handling rare or out-of-distribution (OOD) events expands over time. Conventional static benchmarks and single-shot robustness tests overlook the reality that environments evolve and that models, if left unchallenged, can drift into maladaptation (e.g., reward hacking, over-optimization, or atrophy of broader capabilities). We argue that an antifragile approach -- Rather than striving to rapidly reduce current uncertainties, the emphasis is on leveraging those uncertainties to better prepare for potentially greater, more unpredictable uncertainties in the future -- is pivotal for the long-term reliability of open-ended ML systems. In this position paper, we first identify key limitations of static testing, including scenario diversity, reward hacking, and over-alignment. We then explore the potential of antifragile solutions to manage rare events. Crucially, we advocate for a fundamental recalibration of the methods used to measure, benchmark, and continually improve AI safety over the long term, complementing existing robustness approaches by providing ethical and practical guidelines towards fostering an antifragile AI safety community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。