用虚拟环境模拟未来,让语言模型更懂机器人操作
Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
- 用数字孪生模拟物理世界,生成可执行的运动轨迹
- 通过未来视觉反馈优化语言指令理解,提升任务成功率
- 适合需要复杂交互的开放世界机器人控制场景
近期开放世界机器人操作的进步主要依赖视觉语言模型(VLMs)。尽管这些模型在高层规划中表现出强泛化能力,但因缺乏对物理世界的深入理解,在低层控制预测上表现不佳。为此,我们提出一种结合VLM语义推理能力与真实环境物理驱动的交互式数字孪生的模型预测控制框架。通过构建并仿真数字孪生,该方法生成可行的运动轨迹,模拟相应结果,并将未来观测结果回传给VLM,以评估和选择最符合语言指令的任务结果。为增强预训练VLM对复杂场景的理解能力,我们利用数字孪生灵活的渲染功能,在多种新颖、无遮挡视角下合成场景。我们在多样化的复杂操作任务上验证了该方法,相比基于VLM的语言条件机器人控制基线方法表现更优。
原文摘要 · Abstract (English)
Recent advancements in open-world robot manipulation have been largely driven by vision-language models (VLMs). While these models exhibit strong generalization ability in high-level planning, they struggle to predict low-level robot controls due to limited physical-world understanding. To address this issue, we propose a model predictive control framework for open-world manipulation that combines the semantic reasoning capabilities of VLMs with physically-grounded, interactive digital twins of the real-world environments. By constructing and simulating the digital twins, our approach generates feasible motion trajectories, simulates corresponding outcomes, and prompts the VLM with future observations to evaluate and select the most suitable outcome based on language instructions of the task. To further enhance the capability of pre-trained VLMs in understanding complex scenes for robotic control, we leverage the flexible rendering capabilities of the digital twin to synthesize the scene at various novel, unoccluded viewpoints. We validate our approach on a diverse set of complex manipulation tasks, demonstrating superior performance compared to baseline methods for language-conditioned robotic control using VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。