arXiv:2506.22492cs.CYcs.AI2025-06被引 1

提出下一代AI安全研究框架,聚焦可信赖系统构建。

Report on NSF Workshop on Science of Safe AI

  • 组织跨领域研讨会,凝聚安全AI研究共识
  • 明确需发展理论与工具以保障AI可信赖性
  • 适合关注AI安全、可信系统的研究者参考

机器学习的最新进展,尤其是基础模型的出现,为解决社会问题提供了新的技术路径。然而,当前复杂AI模型的推理过程对用户不透明,且其预测结果缺乏安全保障。因此,要实现AI的潜力,必须解决一个关键科学挑战:如何开发不仅准确高效,而且安全可信的AI系统?这一挑战在自主控制系统和机器人领域尤为突出,也催生了美国国家科学基金会(NSF)的“安全学习赋能系统”(SLES)计划。对于更广泛的AI应用,如用户使用聊天机器人或临床医生接收治疗建议,安全性同样重要,但定义更模糊且依赖具体场景。为此,2025年2月26日,美国宾夕法尼亚大学举办了一场全天研讨会,汇聚了参与NSF SLES项目的研究人员及更广泛领域的AI安全研究者。本报告基于工作组讨论成果,提出了一个新研究议程,旨在发展理论、方法与工具,为下一代AI系统奠定安全基础。

原文摘要 · Abstract (English)

Recent advances in machine learning, particularly the emergence of foundation models, are leading to new opportunities to develop technology-based solutions to societal problems. However, the reasoning and inner workings of today's complex AI models are not transparent to the user, and there are no safety guarantees regarding their predictions. Consequently, to fulfill the promise of AI, we must address the following scientific challenge: how to develop AI-based systems that are not only accurate and performant but also safe and trustworthy? The criticality of safe operation is particularly evident for autonomous systems for control and robotics, and was the catalyst for the Safe Learning Enabled Systems (SLES) program at NSF. For the broader class of AI applications, such as users interacting with chatbots and clinicians receiving treatment recommendations, safety is, while no less important, less well-defined with context-dependent interpretations. This motivated the organization of a day-long workshop, held at University of Pennsylvania on February 26, 2025, to bring together investigators funded by the NSF SLES program with a broader pool of researchers studying AI safety. This report is the result of the discussions in the working groups that addressed different aspects of safety at the workshop. The report articulates a new research agenda focused on developing theory, methods, and tools that will provide the foundations of the next generation of AI-enabled systems.

AI安全可信AI研究议程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。