用视频模型模拟机器人在复杂场景下的表现,验证其泛化与安全性。
Evaluating Gemini Robotics Policies in a Veo World Simulator
- 基于Veo视频模型构建可条件生成的仿真系统,支持多视角一致和场景编辑。
- 在1600+真实评估中准确预测8个机器人策略在正常与异常场景下的性能差异。
- 适用于测试机器人泛化能力、安全边界,适合算法开发者和安全评测人员。
生成式世界模型在模拟视觉运动策略于多样化环境中的交互方面具有巨大潜力。前沿视频模型能够以可扩展且通用的方式生成逼真的观测结果和环境交互。然而,视频模型在机器人领域的应用仍主要局限于分布内评估,即与策略训练或基础视频模型微调所用场景相似的情形。本报告证明,视频模型可用于机器人策略评估的全谱场景:从常规性能评估到分布外(OOD)泛化,以及物理与语义安全性的探测。我们构建了一个基于前沿视频基础模型Veo的生成式评估系统,该系统优化了机器人动作条件生成与多视角一致性,并整合了生成式图像编辑与多视角补全技术,以沿多个泛化轴合成真实世界的多样化场景。系统能保持视频模型的原始能力,准确模拟包含新交互物体、新视觉背景和新干扰物的场景。这种保真度使得系统能准确预测不同策略在正常与分布外条件下的相对性能,确定各泛化轴对策略表现的影响程度,并进行策略红队测试,暴露违反物理或语义安全约束的行为。我们通过1600+次针对八个Gemini Robotics策略检查点和五项双臂操作任务的真实世界评估验证了这些能力。
原文摘要 · Abstract (English)
Generative world models hold significant potential for simulating interactions with visuomotor policies in varied environments. Frontier video models can enable generation of realistic observations and environment interactions in a scalable and general manner. However, the use of video models in robotics has been limited primarily to in-distribution evaluations, i.e., scenarios that are similar to ones used to train the policy or fine-tune the base video model. In this report, we demonstrate that video models can be used for the entire spectrum of policy evaluation use cases in robotics: from assessing nominal performance to out-of-distribution (OOD) generalization, and probing physical and semantic safety. We introduce a generative evaluation system built upon a frontier video foundation model (Veo). The system is optimized to support robot action conditioning and multi-view consistency, while integrating generative image-editing and multi-view completion to synthesize realistic variations of real-world scenes along multiple axes of generalization. We demonstrate that the system preserves the base capabilities of the video model to enable accurate simulation of scenes that have been edited to include novel interaction objects, novel visual backgrounds, and novel distractor objects. This fidelity enables accurately predicting the relative performance of different policies in both nominal and OOD conditions, determining the relative impact of different axes of generalization on policy performance, and performing red teaming of policies to expose behaviors that violate physical or semantic safety constraints. We validate these capabilities through 1600+ real-world evaluations of eight Gemini Robotics policy checkpoints and five tasks for a bimanual manipulator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。