arXiv:2604.19522cs.RO2026-04

用视觉语言模型指导双臂机器人安全操作,让动作更智能更听话。

GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation

论文配图:GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation
图 1 · 摘自论文原文
  • 用VLM-RAG把语义理解转为控制参数,实现上下文感知的物理控制
  • 在仿真和真实机器人上实现接近人类时速度降低60%,动作更安全
  • 适合需要人机协作的智能机器人研发人员参考

双臂移动操作需无缝融合高层语义推理与安全柔顺的物理交互,而端到端模型透明性差,传统控制器缺乏上下文。本文提出GenerativeMPC,一种分层人机系统框架,显式将场景语义理解映射为物理控制参数。系统利用视觉-语言模型结合检索增强生成(VLM-RAG),将视觉与语言上下文转化为具身控制约束,输出全身体型模型预测控制(Whole-Body MPC)的动态速度限制与安全裕度。同时,VLM-RAG调节统一阻抗-导纳控制器的虚拟刚度与阻尼增益,实现人机交互中的上下文感知柔顺性。系统基于经验驱动的向量数据库确保参数一致性,无需重新训练。在MuJoCo、IsaacSim及真实双臂平台上的实验表明,靠近人类时速度降低60%,实现了安全、符合社会规范的导航与操作。该工作通过将大规模认知模型融入可预测的高频物理控制环路,推动了以人为中心的网络物理系统发展。

原文摘要 · Abstract (English)

Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challenge that end-to-end models approach opaquely and classical controllers lack the context to address. This paper presents GenerativeMPC, a hierarchical cyber-physical framework that explicitly bridges semantic scene understanding with physical control parameters for bimanual mobile manipulators. The system utilizes a Vision-Language Model with Retrieval-Augmented Generation (VLM-RAG) to translate visual and linguistic context into grounded control constraints, specifically outputting dynamic velocity limits and safety margins for a Whole-Body Model Predictive Controller (MPC). Simultaneously, the VLM-RAG module modulates virtual stiffness and damping gains for a unified impedance-admittance controller, enabling context-aware compliance during human-robot interaction. Our framework leverages an experience-driven vector database to ensure consistent parameter grounding without retraining. Experimental results in MuJoCo, IsaacSim, and on a physical bimanual platform confirm a 60% speed reduction near humans and safe, socially-aware navigation and manipulation through semantic-to-physical parameter grounding. This work advances the field of human-centric cybernetics by grounding large-scale cognitive models into predictable, high-frequency physical control loops.

机器人控制人机交互视觉语言模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。