用视觉语言模型实现人机协作任务的动态重规划,兼顾语义理解与物理可行性。
Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical Dual-Correction
- 将指令映射为动作目标候选,结合内外部纠错模型验证逻辑与视觉结果。
- 实测在真实场景中完成物体固定成功率66.7%,初始工具选择100%准确。
- 适合关注人机协作、视觉语言模型落地应用的研究者与工程师。
人机协同装配要求机器人在模糊修正指令下生成可执行的动作。视觉语言模型(VLM)具备语义推理能力,但可能选择逻辑不一致的目标或误判执行结果。本文提出一种重规划框架,将人类指令映射为动作目标候选(包括抓取位姿和工具选择),并结合内部纠错模型进行预执行逻辑验证,以及外部纠错模型进行后执行视觉验证。该框架融合VLM推理、6-DoF抓取生成与无碰撞轨迹规划。仿真消融实验表明:内部纠错提升候选有效性,外部纠错在低延迟VLM下可实现恢复,但若视觉验证产生误报则会降低成功率。在上肢人形机器人上的实测结果显示:真实场景中物体固定成功率为66.7%,初始工具选择准确率达100%,修正工具选择成功率为75.0%。结果表明该方法可在空间与语义层面实现交互式重规划,同时揭示视觉状态验证是关键瓶颈。
原文摘要 · Abstract (English)
Human-robot collaborative assembly requires robots to interpret ambiguous corrective instructions while producing physically executable motions. Vision-language models (VLMs) provide semantic reasoning but may select logically inconsistent targets or misjudge execution outcomes. We propose a replanning framework that maps human instructions to Action Target candidates, including grasp poses and tool selections, and combines an Internal Correction Model for pre-execution logical verification with an External Correction Model for post-execution visual verification. The framework integrates VLM reasoning with 6-DoF grasp generation and collision-free trajectory planning. Simulation ablations show configuration-dependent effects: internal correction improves candidate validity, whereas external correction enables recovery for a low-latency VLM but can reduce success when visual verification produces false negatives. Experiments with an upper-body humanoid robot achieved 66.7% success in real-world object fixation, 100% in initial tool selection, and 75.0% in corrective tool selection. These results demonstrate interactive replanning across spatial and semantic collaborative tasks while identifying visual-state verification as a key limitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。