arXiv:2502.04512cs.AI2025-02中稿 · ICML被引 4

开放式AI自主演化带来不可预测风险,需提前防范安全问题。

Safety Must Precede the Deployment of Open-Ended AI

  • 提出开放式AI系统安全挑战的分类框架
  • 指出现有安全机制无法应对自主演化带来的风险
  • 呼吁在大规模部署前协同推进安全研究

AI发展依赖于基础模型与好奇心驱动学习,以提升能力与适应性。其中,开放式AI系统能自主、无限生成新行为或解决方案,引发自演化智能体与长周期探索的关注。本文指出,此类系统因脱离初始设计假设,易产生不可预测性、目标错位及控制失效等独特安全风险,与任务限定或静态模型的风险本质不同。现有安全框架难以应对这些挑战,必须提前介入。论文构建了关键挑战的分类体系,探讨研究机遇,并呼吁跨领域协作,推动开放式AI的安全与负责任发展。

原文摘要 · Abstract (English)

AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability and adaptability. Within this landscape, open-endedness, where AI agents autonomously and indefinitely generate novel behaviors, representations, or solutions, has gained increasing interest. This has become relevant in the context of self-evolving agents and long-horizon discovery. This position paper argues that the defining properties of open-ended AI systems introduce a distinct and underexplored class of safety challenges, including loss of predictability, emergent misalignment, and difficulties in maintaining effective control as systems evolve beyond their initial design assumptions, that must be addressed preemptively. These challenges differ qualitatively from those associated with task-bounded or static models and are unlikely to be addressed by existing safety frameworks alone, which is why these risks must be examined proactively, before large-scale deployment. The paper proposes a taxonomy for key challenges, discusses research opportunities, and calls for coordinated action to support the safe and responsible development of open-ended AI.

AI安全开放式系统风险防控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。