让机器人在工业场景中精准操作,靠自然手势+文本指令自动避障。
From Perception to Assistance: Open-Vocabulary Shared Autonomy for Robotic Manipulation

- 用视觉语言模型理解文本指令,结合手势和摄像头追踪目标。
- 控制算法实时避障,保持机械臂离障碍物至少18厘米。
- 支持手势触发自主抓取,适合工业复杂环境远程操控。
在工业环境中遥控机械臂需要高精度,仅靠摄像头难以实现。本文提出一种共享自主框架,利用单个RGB-D相机捕捉操作者手臂动作与手势,无需穿戴设备或校准。通过自由文本提示,由视觉-语言模型在机械臂摄像头中定位目标,并由可提示视频分割模型在多视角下持续跟踪,生成始终避开障碍物的抓取位姿。每个指令由基于GPU加速的模型预测控制器执行,实时避免自碰撞与环境碰撞,基于势场将操作者指令修正至目标点。可通过手势触发自主模式完成抓取,无需额外感知流程。在四足移动机械臂上验证:相对于动作捕捉真值,位置均方根误差为59毫米;即便操作者刻意命令接近障碍物6厘米,仍能保持至少18厘米距离。在阀门操作与拾放任务中,完整框架全部成功,去除任一模块均导致互补性失败,自主执行每项任务在五次中成功四次。
原文摘要 · Abstract (English)
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces alone struggle to deliver. The operator must align the end-effector with a target in clutter, under limited depth perception, and without colliding with the surrounding structures. This paper presents a shared-autonomy framework that assists the operator throughout this process. A single RGB-D camera captures the operator's arm motion and hand gestures without wearables, fiducials, or a calibration stage. The intended target is specified by a free-form text prompt, grounded by a vision-language model in the robot's gripper camera, and tracked across its onboard cameras by a promptable video-segmentation model, resulting in a grasp frame continuously separated from the obstacle map. Every commanded motion is executed by a GPU-accelerated model-predictive controller that enforces self- and environment-collision avoidance against an online volumetric reconstruction, while a potential field corrects the operator's reference toward the grounded target during the final approach. An autonomous mode can be gesture-triggered to complete the grasp on the same target without a separate perception pipeline. The framework is validated on a quadruped mobile manipulator. The interface achieves a positional RMSE of 59 mm relative to motion-capture ground truth, and the controller keeps the arm at least 18 cm from obstacles while the operator deliberately commands the arm into them by 6 cm. In an industrial valve manipulation and a pick-and-place task, the full framework succeeded in all trials, while ablating either the collision or the assistance module produced failures through complementary mechanisms, and autonomous execution succeeded in four of five trials per task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。