arXiv:2602.24235cs.ROcs.AI2026-02被引 1

让机器人任务规划更安全且能适应新安全规则。

SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems

  • 用两阶段训练提升LLM在任务规划中的安全性与泛化能力。
  • 在多领域测试中,对新安全规则的适应性显著优于主流基线。
  • 适合关注机器人安全、智能规划的开发者与研究者。

机器人系统中的安全关键型任务规划仍具挑战:传统规划器扩展性差,基于强化学习的方法泛化能力弱,而基础大语言模型无法保证安全。为此,我们提出可泛化安全的大语言模型SafeGen-LLM,不仅能提升任务计划的安全性,还能在多种领域中泛化到新的安全属性。我们首先构建了包含显式安全约束的多领域规划域定义语言3(PDDL3)基准。随后引入两阶段后训练框架:在符合约束的规划数据集上进行监督微调(SFT),以学习规划语法与语义;再通过由形式验证生成的细粒度奖励机引导的组相对策略优化(GRPO),结合课程学习,实现安全对齐并更好处理复杂任务。大量实验表明,SafeGen-LLM在多领域规划任务和多种输入格式(如PDDL与自然语言)下均展现出强安全泛化能力,优于前沿专有基线。

原文摘要 · Abstract (English)

Safety-critical task planning in robotic systems remains challenging: classical planners suffer from poor scalability, Reinforcement Learning (RL)-based methods generalize poorly, and base Large Language Models (LLMs) cannot guarantee safety. To address this gap, we propose safety-generalizable large language models, named SafeGen-LLM. SafeGen-LLM can not only enhance the safety satisfaction of task plans but also generalize well to novel safety properties in various domains. We first construct a multi-domain Planning Domain Definition Language 3 (PDDL3) benchmark with explicit safety constraints. Then, we introduce a two-stage post-training framework: Supervised Fine-Tuning (SFT) on a constraint-compliant planning dataset to learn planning syntax and semantics, and Group Relative Policy Optimization (GRPO) guided by fine-grained reward machines derived from formal verification to enforce safety alignment and by curriculum learning to better handle complex tasks. Extensive experiments show that SafeGen-LLM achieves strong safety generalization and outperforms frontier proprietary baselines across multi-domain planning tasks and multiple input formats (e.g., PDDLs and natural language).

机器人规划大模型安全强化学习形式验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。