为机器人生成安全行为准则,提升语义安全能力。
Generating Robot Constitutions & Benchmarks for Semantic Safety
- 用生成式AI从真实场景和伤情报告中构建危险情境数据集。
- 自动生成可自动更新的行为准则,对齐人类偏好,准确率超84%。
- 适合关注机器人安全、伦理对齐的研究者与开发者。
长期以来,机器人安全研究主要聚焦于碰撞规避和局部危险降低。随着大视觉语言模型(VLMs)的发展,机器人具备了更高层次的语义理解与人机自然语言交互能力。然而,这些模型存在幻觉或越狱等已知漏洞,却已被赋予操控可与现实世界物理交互的机器人。这可能导致危险行为,使语义安全成为紧迫问题。本文贡献有二:其一,为应对新兴风险,我们发布ASIMOV基准,一个大规模、全面的数据集,用于评估和提升作为机器人‘大脑’的基础模型的语义安全性。我们的数据生成方法高度可扩展:通过文本与图像生成技术,从真实视觉场景和医院伤情报告中生成不良情境。其二,我们提出框架,从真实数据自动生成机器人行为准则,利用宪法AI机制引导机器人行为。设计新颖的自动修订流程,能引入行为规则的细微差别,提升与人类偏好在行为可接受性与安全性上的对齐度。我们在不同长度的多种准则间探索通用性与特异性权衡,证明机器人能有效拒绝非宪法行为。使用生成准则在ASIMOV基准上实现最高84.3%的对齐率,优于无准则基线与人工编写准则。数据可在asimov-benchmark.github.io获取。
原文摘要 · Abstract (English)
Until recently, robotics safety research was predominantly about collision avoidance and hazard reduction in the immediate vicinity of a robot. Since the advent of large vision and language models (VLMs), robots are now also capable of higher-level semantic scene understanding and natural language interactions with humans. Despite their known vulnerabilities (e.g. hallucinations or jail-breaking), VLMs are being handed control of robots capable of physical contact with the real world. This can lead to dangerous behaviors, making semantic safety for robots a matter of immediate concern. Our contributions in this paper are two fold: first, to address these emerging risks, we release the ASIMOV Benchmark, a large-scale and comprehensive collection of datasets for evaluating and improving semantic safety of foundation models serving as robot brains. Our data generation recipe is highly scalable: by leveraging text and image generation techniques, we generate undesirable situations from real-world visual scenes and human injury reports from hospitals. Secondly, we develop a framework to automatically generate robot constitutions from real-world data to steer a robot's behavior using Constitutional AI mechanisms. We propose a novel auto-amending process that is able to introduce nuances in written rules of behavior; this can lead to increased alignment with human preferences on behavior desirability and safety. We explore trade-offs between generality and specificity across a diverse set of constitutions of different lengths, and demonstrate that a robot is able to effectively reject unconstitutional actions. We measure a top alignment rate of 84.3% on the ASIMOV Benchmark using generated constitutions, outperforming no-constitution baselines and human-written constitutions. Data is available at asimov-benchmark.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。