分析前沿AI五大风险,提出可落地的安全应对方案。
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5
- 从五个维度评估大模型与代理型AI的潜在威胁。
- 发现新模型存在链式说服、自主演进等失控风险。
- 适合关注AI安全与治理的研究者与政策制定者。
为理解并识别快速发展的先进人工智能(AI)模型带来的前所未有风险,本报告对前沿风险进行了全面评估。随着大型语言模型(LLMs)通用能力迅速提升及代理型AI广泛普及,本版技术报告对五大关键维度——网络攻击、操纵与影响、战略欺骗、不受控的AI研发、自我复制——进行了更新和细化分析。具体包括:引入更复杂的网络攻击场景;评估新发布大模型间通过大模型进行链式说服的风险;新增关于涌现性错位的实验;聚焦代理在自主扩展记忆与工具集过程中产生的‘误演化’问题;同时监控OpenClaw在Moltbook交互中的安全表现。针对自我复制,引入资源受限的新情景。更重要的是,本文提出并验证了一系列稳健缓解策略,为前沿AI的安全部署提供初步的技术路径与行动指南。该工作反映了当前对前沿AI风险的理解,并呼吁集体行动以应对这些挑战。
原文摘要 · Abstract (English)
To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, Frontier AI Risk Management Framework in Practice presents a comprehensive assessment of their frontier risks. As Large Language Models (LLMs) general capabilities rapidly evolve and the proliferation of agentic AI, this version of the risk analysis technical report presents an updated and granular assessment of five critical dimensions: cyber offense, persuasion and manipulation, strategic deception, uncontrolled AI R\&D, and self-replication. Specifically, we introduce more complex scenarios for cyber offense. For persuasion and manipulation, we evaluate the risk of LLM-to-LLM persuasion on newly released LLMs. For strategic deception and scheming, we add the new experiment with respect to emergent misalignment. For uncontrolled AI R\&D, we focus on the ``mis-evolution'' of agents as they autonomously expand their memory substrates and toolsets. Besides, we also monitor and evaluate the safety performance of OpenClaw during the interaction on the Moltbook. For self-replication, we introduce a new resource-constrained scenario. More importantly, we propose and validate a series of robust mitigation strategies to address these emerging threats, providing a preliminary technical and actionable pathway for the secure deployment of frontier AI. This work reflects our current understanding of AI frontier risks and urges collective action to mitigate these challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。