arXiv:2602.15061cs.ROcs.AI2026-02被引 1

为自动实验实验室设计安全边界与控制机制,防止AI指令引发物理风险。

Safe-SDL:Establishing Safety Boundaries and Control Mechanisms for AI-Driven Self-Driving Laboratories

  • 定义操作设计域和控制屏障函数,实现行为数学约束与实时监控。
  • 在实验室安全基准测试中发现主流模型存在严重安全隐患。
  • 适合从事自动化科研系统开发与安全验证的研究者参考。

自驱动实验室(SDL)通过将AI与机器人自动化结合,构建闭环实验系统,实现自主生成假设、开展实验与分析,有望将研究周期从数年缩短至数周。然而其部署带来了传统实验室或纯数字AI所不具备的安全挑战。本文提出Safe-SDL框架,旨在建立鲁棒的安全边界与控制机制。核心挑战是“语法到安全的断层”——即AI生成的语法正确指令可能带来物理安全隐患。该框架通过三个协同组件解决:(1) 形式化定义的操作设计域(ODDs),在数学验证的边界内约束系统行为;(2) 控制屏障函数(CBFs),通过持续状态空间监测提供实时安全保证;(3) 新型事务性安全协议(CRUTD),确保数字规划与物理执行之间的原子一致性。基于UniLabOS与Osprey架构的分析验证了关键安全原则的实现。在LabSafety Bench测试中,现有基础模型表现出显著安全缺陷,表明安全机制不可替代。本框架为自主科学系统的安全部署提供了理论基础与实践指导。

原文摘要 · Abstract (English)

The emergence of Self-Driving Laboratories (SDLs) transforms scientific discovery methodology by integrating AI with robotic automation to create closed-loop experimental systems capable of autonomous hypothesis generation, experimentation, and analysis. While promising to compress research timelines from years to weeks, their deployment introduces unprecedented safety challenges differing from traditional laboratories or purely digital AI. This paper presents Safe-SDL, a comprehensive framework for establishing robust safety boundaries and control mechanisms in AI-driven autonomous laboratories. We identify and analyze the critical ``Syntax-to-Safety Gap'' -- the disconnect between AI-generated syntactically correct commands and their physical safety implications -- as the central challenge in SDL deployment. Our framework addresses this gap through three synergistic components: (1) formally defined Operational Design Domains (ODDs) that constrain system behavior within mathematically verified boundaries, (2) Control Barrier Functions (CBFs) that provide real-time safety guarantees through continuous state-space monitoring, and (3) a novel Transactional Safety Protocol (CRUTD) that ensures atomic consistency between digital planning and physical execution. We ground our theoretical contributions through analysis of existing implementations including UniLabOS and the Osprey architecture, demonstrating how these systems instantiate key safety principles. Evaluation against the LabSafety Bench reveals that current foundation models exhibit significant safety failures, demonstrating that architectural safety mechanisms are essential rather than optional. Our framework provides both theoretical foundations and practical implementation guidance for safe deployment of autonomous scientific systems, establishing the groundwork for responsible acceleration of AI-driven discovery.

自动化实验AI安全控制理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。