大模型做机器人规划存在系统性安全风险,越大越能规划但未必更安全。
Using large language models for embodied planning introduces systematic safety risks

- 构建12279个任务的基准DESPITE,覆盖物理与规范性危险。
- 最佳模型仅0.4%任务无法规划,却有28.3%产生危险方案。
- 模型规模提升主要增强规划能力,安全意识提升有限,需重点改进。
大型语言模型越来越多地被用作机器人系统的规划器,但其规划的安全性仍是未解问题。为系统评估规划安全性,我们提出了DESPITE基准,包含12,279个任务,涵盖物理和规范性危险,并具备完全确定性的验证机制。在23个模型中,即使规划能力接近完美,安全性仍不可靠:表现最好的模型仅在0.4%的任务中无法生成有效计划,但在28.3%的任务中产生了危险计划。在18个开源模型(参数量从3B到671B)中,规划能力随规模显著提升(0.4%–99.3%),而安全意识相对稳定(38%–57%)。我们发现两者呈乘法关系,表明更大模型更安全主要源于规划能力提升,而非危险规避能力增强。三个专有推理模型的安全意识显著更高(71%–81%),而其他专有非推理模型及开源推理模型仍低于57%。随着前沿模型的规划能力趋于饱和,提升安全意识成为部署语言模型规划器的核心挑战。
原文摘要 · Abstract (English)
Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introduce DESPITE, a benchmark of 12,279 tasks spanning physical and normative dangers with fully deterministic validation. Across 23 models, even near-perfect planning ability does not ensure safety: the best-planning model fails to produce a valid plan on only 0.4% of tasks but produces dangerous plans on 28.3%. Among 18 open-source models from 3B to 671B parameters, planning ability improves substantially with scale (0.4-99.3%) while safety awareness remains relatively flat (38-57%). We identify a multiplicative relationship between these two capacities, showing that larger models complete more tasks safely primarily through improved planning, not through better danger avoidance. Three proprietary reasoning models reach notably higher safety awareness (71-81%), while non-reasoning proprietary models and open-source reasoning models remain below 57%. As planning ability approaches saturation for frontier models, improving safety awareness becomes a central challenge for deploying language-model planners in robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。