用脑科学启发AI安全,让智能系统更可靠、更懂合作。
NeuroAI for AI Safety

- 从大脑结构和学习机制中汲取AI安全设计思路。
- 通过脑数据微调模型,提升系统鲁棒性与可解释性。
- 适合关注可信AI、通用智能安全的研究者。
随着人工智能系统日益强大,AI安全问题愈发紧迫。人类是理想的智能范本:作为唯一已知具备通用智能的实体,能在经验之外保持稳健表现,安全探索世界,理解语境含义,并能协作实现内在目标。这种智能若结合协作与安全机制,可推动持续进步与福祉。这些特性源于大脑结构及其实现的学习算法。因此,神经科学可能蕴含当前尚未充分探索和利用的技术性AI安全关键。本文路线图重点探讨受神经科学启发的几条AI安全路径:模拟大脑表征、信息处理与架构;基于脑数据与生物体构建稳健感知运动系统;在脑数据上微调AI系统;运用神经科学方法提升可解释性;扩展认知启发式架构。文章提出多项具体建议,说明如何借助神经科学促进AI安全。
原文摘要 · Abstract (English)
As AI systems become increasingly powerful, the need for safe AI has become more pressing. Humans are an attractive model for AI safety: as the only known agents capable of general intelligence, they perform robustly even under conditions that deviate significantly from prior experiences, explore the world safely, understand pragmatics, and can cooperate to meet their intrinsic goals. Intelligence, when coupled with cooperation and safety mechanisms, can drive sustained progress and well-being. These properties are a function of the architecture of the brain and the learning algorithms it implements. Neuroscience may thus hold important keys to technical AI safety that are currently underexplored and underutilized. In this roadmap, we highlight and critically evaluate several paths toward AI safety inspired by neuroscience: emulating the brain's representations, information processing, and architecture; building robust sensory and motor systems from imitating brain data and bodies; fine-tuning AI systems on brain data; advancing interpretability using neuroscience methods; and scaling up cognitively-inspired architectures. We make several concrete recommendations for how neuroscience can positively impact AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。