arXiv:2512.04308cs.RO2025-12被引 3

构建多模态大模型评估基准,推动机器人安全操作能力发展

ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models

  • 设计23个分阶段任务,覆盖电、化学及人因风险
  • 集成感知-推理-规划全流程,支持多模态动作表示
  • 提供可复现实验与安全成功率等标准化指标

近期多模态大模型的发展为具身智能,特别是机器人操作带来了新机遇。这些模型在泛化和推理方面展现出强大潜力,但在真实场景中实现可靠且负责任的机器人行为仍是开放挑战。在高风险环境中,机器人不仅需完成基础任务,还需具备风险感知、道德决策与物理可行规划能力。我们提出ResponsibleRobotBench,一个从仿真到现实系统化的基准测试平台,包含23个多阶段任务,涵盖电气、化学及人因危害等多种风险类型,并具有不同物理与规划复杂度。任务要求智能体检测并缓解风险,进行安全推理、规划动作序列,并在必要时请求人类协助。该基准提供通用评估框架,支持多种动作表示的多模态模型,整合视觉感知、上下文学习、提示构造、危险识别、推理规划与物理执行。同时提供丰富多模态数据集,支持可复现实验,包含成功率、安全率、安全成功率等标准化指标。通过大量实验设置,该基准可分析不同风险类别、任务类型与代理配置的表现。通过强调物理可靠性、泛化能力和决策安全性,该基准为发展可信、真实的灵巧机器人系统奠定基础。

原文摘要 · Abstract (English)

Recent advances in large multimodal models have enabled new opportunities in embodied AI, particularly in robotic manipulation. These models have shown strong potential in generalization and reasoning, but achieving reliable and responsible robotic behavior in real-world settings remains an open challenge. In high-stakes environments, robotic agents must go beyond basic task execution to perform risk-aware reasoning, moral decision-making, and physically grounded planning. We introduce ResponsibleRobotBench, a systematic benchmark designed to evaluate and accelerate progress in responsible robotic manipulation from simulation to real world. This benchmark consists of 23 multi-stage tasks spanning diverse risk types, including electrical, chemical, and human-related hazards, and varying levels of physical and planning complexity. These tasks require agents to detect and mitigate risks, reason about safety, plan sequences of actions, and engage human assistance when necessary. Our benchmark includes a general-purpose evaluation framework that supports multimodal model-based agents with various action representation modalities. The framework integrates visual perception, context learning, prompt construction, hazard detection, reasoning and planning, and physical execution. It also provides a rich multimodal dataset, supports reproducible experiments, and includes standardized metrics such as success rate, safety rate, and safe success rate. Through extensive experimental setups, ResponsibleRobotBench enables analysis across risk categories, task types, and agent configurations. By emphasizing physical reliability, generalization, and safety in decision-making, this benchmark provides a foundation for advancing the development of trustworthy, real-world responsible dexterous robotic systems. https://sites.google.com/view/responsible-robotbench

机器人安全多模态模型具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。