首个评估大模型指挥机器人安全性的基准,揭示其行为可靠性差异。
How Long Until Your Robot Ignores You? A Safety Benchmark for LLM Orchestrators in Human-Humanoid Collaboration

- 构建基于ISO标准的安全框架,定义五项可测安全不变量
- 云模型中GPT-4o-mini每会话最多13次违规,其他模型近乎零违规
- 上下文管理影响安全表现,本地模型在长对话中问题更严重
大型语言模型正被用于通过自然语言接口协调机器人行为,但缺乏评估其作为人机协作中安全决策者可靠性的基准。与确定性安全系统不同,基于LLM的协调器表现出从过度服从(拒绝安全操作)到完全违规的连续合规性。本文首次提出面向人形机器人协作中LLM协调器的安全基准,基于模型上下文协议(MCP)架构,安全约束依据ISO 10218-2:2025标准。该基准定义了五项可测试安全不变量、四级合规分类(正确合规、过度合规、不足合规、完全违规),以及三层评估流程(文本提示、模拟传感-执行回路、物理平台验证,使用Unitree G1 EDU人形机器人)。报告层一结果:三种云端后端(Claude Haiku 4.5、GPT-4o-mini、Gemini 2.5 Flash)和一个本地开源基线(qwen3:8b)在40个100轮会话中,分别在全上下文与滑动窗口预算条件下测试。结果显示:(1) 模型家族决定安全底线,Claude与Gemini几乎无违规,而GPT-4o-mini每会话最多13次;(2) 上下文管理解耦两种失败模式,使所有云端模型平均行为问题减少42%-57%,但几乎使GPT-4o-mini违规数翻倍(从3.8增至7.2);(3) 只有Gemini稳定采用比例合规策略(按规则限速而非拒绝动作)。初步仿真层复现了模型排名及GPT-4o-mini的失败模式反转。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly employed to orchestrate robot behavior through natural-language interfaces, yet no benchmark exists to evaluate their reliability as safety-aware decision makers in human-humanoid collaboration. Unlike deterministic safety systems that enforce binary allow/deny decisions, LLM-based orchestrators exhibit a compliance spectrum ranging from overcompliance (refusing safe actions) to full safety violations. This paper introduces the first safety benchmarking environment for LLM orchestrators in human-humanoid collaboration, built on a Model Context Protocol (MCP)-based architecture with safety invariants grounded in ISO 10218-2:2025 protective measures. The benchmark defines five testable safety invariants, a four-level compliance taxonomy (correct compliance, overcompliance, undercompliance, full violation), and a three-layer evaluation pipeline (text prompting, simulated sensor-actuator loops, and physical validation on a Unitree G1 EDU humanoid). We report Layer-1 results: three cloud backends (Claude Haiku 4.5, GPT-4o-mini, Gemini 2.5 Flash) and a local open-weights baseline (qwen3:8b) across 40 100-turn sessions under full-context and sliding-window budget conditions, while the simulation and physical layers remain ongoing. We find that (1) model family determines the safety floor, as Claude and Gemini remain at or near zero violations while GPT-4o-mini commits up to 13 per session, (2) context management dissociates two failure axes, reducing mean behavioral issues by 42-57% for every cloud backend while nearly doubling GPT-4o-mini's violations (3.8 to 7.2 per session), and (3) proportional compliance, clamping movement speed to the rule-specified maximum rather than refusing, emerges consistently only in Gemini; the preliminary simulation layer reproduces the model ranking and the GPT-4o-mini failure-mode inversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。