用AI生成难指令测试机器人模型安全与性能
Embodied Red Teaming for Auditing Robotic Foundation Models
- 用视觉语言模型自动生成挑战性指令
- 顶尖机器人模型在新指令下失败率超50%
- 适合评估机器人安全性的研究者使用
语言驱动的机器人模型有望根据自然语言指令完成多种任务,但其安全性和有效性评估仍具挑战,因任务表述方式多样,难以全面测试。现有基准存在两大局限:依赖有限的人类生成指令,遗漏大量复杂场景;仅关注任务完成度,忽视安全性如避免损坏。为此,本文提出具身红队测试(Embodied Red Teaming, ERT),利用视觉语言模型(VLM)自动化生成情境相关的高难度指令,以全面检验模型表现。实验表明,当前最先进的语言条件机器人模型在ERT生成的指令下,超过50%出现失败或不安全行为,凸显现有基准在真实场景评估中的不足。代码与视频见:https://s-karnik.github.io/embodied-red-team-project-page。
原文摘要 · Abstract (English)
Language-conditioned robot models have the potential to enable robots to perform a wide range of tasks based on natural language instructions. However, assessing their safety and effectiveness remains challenging because it is difficult to test all the different ways a single task can be phrased. Current benchmarks have two key limitations: they rely on a limited set of human-generated instructions, missing many challenging cases, and focus only on task performance without assessing safety, such as avoiding damage. To address these gaps, we introduce Embodied Red Teaming (ERT), a new evaluation method that generates diverse and challenging instructions to test these models. ERT uses automated red teaming techniques with Vision Language Models (VLMs) to create contextually grounded, difficult instructions. Experimental results show that state-of-the-art language-conditioned robot models fail or behave unsafely on ERT-generated instructions, underscoring the shortcomings of current benchmarks in evaluating real-world performance and safety. Code and videos are available at: https://s-karnik.github.io/embodied-red-team-project-page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。