安全不是看当前行为,而是看系统能否持续被纠正。
Agentic Safety is an Epistemic Property, Not a Behavioral One
- 把安全视为系统可被未来纠正的能力,而非仅看当前表现
- 提出‘可教性’概念:在有限干预下仍能保持修正空间
- 适合研究通用AI长期安全与自我改进系统的学者
当前AI安全涵盖预训练干预、后训练对齐、部署控制、监控与红队测试,这些方法虽必要,但主要验证系统当前行为的快照。随着AI系统变得更强大、动态、具身化且具备自我改进能力,仅关注当前行为已不充分:安全不仅关乎当下是否合规,更在于系统在学习、适应、行动和自我修改过程中,能否持续保持可纠正性。本文主张,安全应被视为演化学习者的一种认知属性,而非仅是当前策略的行为属性。我们引入‘可教性’(teachability)概念,指在人类、制度或环境有限干预下,仍能维持未来矫正杠杆的能力。我们指出,先进系统可能表面仍具胜任力,却已丧失未来修正所需的表征、算法或元决策条件。因此,安全的先进AI不仅要当前表现良好,还必须在未来仍可被纠正。
原文摘要 · Abstract (English)
Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they primarily certify snapshots of system behavior. As AI systems become more capable, dynamic, embodied, and self-improving, this snapshot view becomes incomplete: safety depends not only on whether a system behaves acceptably now, but whether it remains correctable as it learns, adapts, acts, and modifies itself over time. This paper argues that safety should therefore be treated as an epistemic property of the evolving learner, not merely a behavioral property of the current policy. We introduce teachability as the capacity to preserve future corrective leverage under bounded human, institutional, or environmental intervention. We argue that advanced systems can retain visible competence while eroding the representational, algorithmic, or meta-decision conditions needed for future correction. Safe advanced AI systems must not only behave acceptably now; they must remain teachable later.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。