arXiv:2511.23300cs.RO2025-11被引 1

用视觉语言模型让机器人根据场景自动调节动作柔韧度和速度,更安全。

SafeHumanoid: VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot

  • 通过视觉语言模型+检索生成,理解场景并调度机器人阻抗参数
  • 在有人/无人场景中自适应调整刚度、阻尼和速度,任务成功率超95%
  • 适合需要人机协作的工业或服务场景,尤其关注安全合规

安全可信的人机交互要求机器人不仅完成任务,还需根据环境上下文和人体接近程度调节阻抗与速度。我们提出SafeHumanoid,一种基于第一视角视觉的控制管道,将视觉语言模型(VLM)与检索增强生成(RAG)结合,用于调度人形机器人的阻抗与速度参数。第一视角图像通过结构化VLM提示处理,嵌入后与经验证的场景数据库匹配,并通过逆运动学映射为关节级阻抗指令。我们在有无人类参与的桌面操作任务上进行评估,包括擦拭、物品交接和液体倾倒。结果表明,该系统能上下文感知地调整刚度、阻尼和速度曲线,在保持任务成功的同时提升安全性。尽管当前推理延迟高达1.4秒,限制了高度动态场景下的响应能力,但实验验证了语义驱动阻抗控制是实现更安全、符合标准的人形机器人协作的可行路径。

原文摘要 · Abstract (English)

Safe and trustworthy Human Robot Interaction (HRI) requires robots not only to complete tasks but also to regulate impedance and speed according to scene context and human proximity. We present SafeHumanoid, an egocentric vision pipeline that links Vision Language Models (VLMs) with Retrieval-Augmented Generation (RAG) to schedule impedance and velocity parameters for a humanoid robot. Egocentric frames are processed by a structured VLM prompt, embedded and matched against a curated database of validated scenarios, and mapped to joint-level impedance commands via inverse kinematics. We evaluate the system on tabletop manipulation tasks with and without human presence, including wiping, object handovers, and liquid pouring. The results show that the pipeline adapts stiffness, damping, and speed profiles in a context-aware manner, maintaining task success while improving safety. Although current inference latency (up to 1.4 s) limits responsiveness in highly dynamic settings, SafeHumanoid demonstrates that semantic grounding of impedance control is a viable path toward safer, standard-compliant humanoid collaboration.

人机交互阻抗控制视觉语言模型机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。