用生成图像模拟环境变化,提前预测机器人操作失败
Predictive Red Teaming: Breaking Policies Without Breaking Robots
- 通过生成图像修改环境因素,自动测试视觉动作策略的脆弱点
- 预测成功率与真实结果误差小于0.19,500+次硬件实验验证
- 可指导针对性数据收集,提升基线性能2-7倍,适合强化训练者
基于模仿学习训练的视觉运动策略虽能完成复杂操作任务,但对光照、视觉干扰和物体位置等环境变化极为敏感。这些漏洞往往难以预判,且暴露需耗时昂贵的实物测试。本文提出预测性红队测试:在不进行硬件评估的前提下,预测策略在非标准场景下的性能退化。为此,我们开发了RoboART——自动化红队测试流程,通过生成式图像编辑改变观测条件,并利用策略特异性异常检测器预测每种变化下的表现。在12类非标准条件下,对500+次硬件试验的视觉运动扩散策略进行验证,结果显示预测成功率与真实成功率平均差异低于0.19。此外,预测出的高风险场景可用于指导数据采集,针对性微调使基线性能提升2-7倍。
原文摘要 · Abstract (English)
Visuomotor policies trained via imitation learning are capable of performing challenging manipulation tasks, but are often extremely brittle to lighting, visual distractors, and object locations. These vulnerabilities can depend unpredictably on the specifics of training, and are challenging to expose without time-consuming and expensive hardware evaluations. We propose the problem of predictive red teaming: discovering vulnerabilities of a policy with respect to environmental factors, and predicting the corresponding performance degradation without hardware evaluations in off-nominal scenarios. In order to achieve this, we develop RoboART: an automated red teaming (ART) pipeline that (1) modifies nominal observations using generative image editing to vary different environmental factors, and (2) predicts performance under each variation using a policy-specific anomaly detector executed on edited observations. Experiments across 500+ hardware trials in twelve off-nominal conditions for visuomotor diffusion policies demonstrate that RoboART predicts performance degradation with high accuracy (less than 0.19 average difference between predicted and real success rates). We also demonstrate how predictive red teaming enables targeted data collection: fine-tuning with data collected under conditions predicted to be adverse boosts baseline performance by 2-7x.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。