arXiv:2601.14874cs.RO2026-01中稿 · publication at LBR…被引 2

用视觉语言模型自动选机器人操控参数,让仿人机器人更灵活应对不同任务。

HumanoidVLM: Vision-Language-Guided Impedance Control for Contact-Rich Humanoid Manipulation

  • 通过视觉语言模型+检索生成,从图像直接获取适配的阻抗参数和抓取角度。
  • 在14种场景下检索准确率达93%,真实实验中位置误差保持在1-3.5厘米内。
  • 适合需要自适应接触操作的仿人机器人研究者,系统可解释性强。

仿人机器人需针对不同物体和任务调整接触行为,但现有控制器多依赖固定、人工调校的阻抗增益与夹爪设置。本文提出HumanoidVLM,一种基于视觉-语言的检索框架,使Unitree G1仿人机器人能直接从自身视角的RGB图像中选取合适的笛卡尔阻抗参数与夹爪配置。系统结合视觉-语言模型进行语义任务识别,并通过基于FAISS的检索增强生成(RAG)模块,从两个自建数据库中检索出经实验验证的刚度-阻尼组合及对象特异性抓取角度,再由任务空间阻抗控制器执行柔性操作。我们在14个视觉场景上评估该系统,检索准确率达到93%。真实世界实验显示稳定交互动态,垂直方向跟踪误差通常在1-3.5厘米之间,虚拟力与任务相关的阻抗设定一致。结果表明,将语义感知与基于检索的控制相结合,是实现自适应仿人操作的一种可解释路径。

原文摘要 · Abstract (English)

Humanoid robots must adapt their contact behavior to diverse objects and tasks, yet most controllers rely on fixed, hand-tuned impedance gains and gripper settings. This paper introduces HumanoidVLM, a vision-language driven retrieval framework that enables the Unitree G1 humanoid to select task-appropriate Cartesian impedance parameters and gripper configurations directly from an egocentric RGB image. The system couples a vision-language model for semantic task inference with a FAISS-based Retrieval-Augmented Generation (RAG) module that retrieves experimentally validated stiffness-damping pairs and object-specific grasp angles from two custom databases, and executes them through a task-space impedance controller for compliant manipulation. We evaluate HumanoidVLM on 14 visual scenarios and achieve a retrieval accuracy of 93%. Real-world experiments show stable interaction dynamics, with z-axis tracking errors typically within 1-3.5 cm and virtual forces consistent with task-dependent impedance settings. These results demonstrate the feasibility of linking semantic perception with retrieval-based control as an interpretable path toward adaptive humanoid manipulation.

仿人机器人视觉语言阻抗控制检索生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。