用视觉语言模型提升人机协作装配的感知与规划鲁棒性
From Perception to Symbolic Task Planning: Vision-Language Guided Human-Robot Collaborative Structured Assembly
- 通过视觉语言模型将摄像头数据与设计规范对齐,生成可验证的装配状态
- 97%的状态合成准确率,在人机交互下仍能保持任务推进可行性
- 仅在状态偏差时最小化重规划,适合动态复杂装配场景
结构化装配中的人机协作需要在感知噪声和人为干预下实现可靠的状态估计与自适应任务规划。为此,我们提出一个面向人机协同装配的设计驱动型感知-规划框架,包含两个耦合模块:模块一(感知到符号状态,PSS)利用基于视觉语言模型的智能体,将RGB-D观测与设计规范及领域知识对齐,生成可验证的符号化装配状态,并输出已安装与未安装组件集合用于在线状态追踪;模块二(人感知规划与重规划,HPR)执行任务级多机器人分配,仅当实际状态偏离预期时才更新计划,采用最小变更重规划策略,以在人介入下保持计划稳定性。我们在27个构件的木结构框架装配任务上验证了该框架,PSS模块达到97%的状态合成准确率,HPR模块在多种人机协作场景中维持了可行的任务进展。结果表明,将视觉语言模型感知与知识驱动规划结合,显著提升了动态条件下的状态估计与任务规划鲁棒性。
原文摘要 · Abstract (English)
Human-robot collaboration (HRC) in structured assembly requires reliable state estimation and adaptive task planning under noisy perception and human interventions. To address these challenges, we introduce a design-grounded human-aware planning framework for human-robot collaborative structured assembly. The framework comprises two coupled modules. Module I, Perception-to-Symbolic State (PSS), employs vision-language models (VLMs) based agents to align RGB-D observations with design specifications and domain knowledge, synthesizing verifiable symbolic assembly states. It outputs validated installed and uninstalled component sets for online state tracking. Module II, Human-Aware Planning and Replanning (HPR), performs task-level multi-robot assignment and updates the plan only when the observed state deviates from the expected execution outcome. It applies a minimal-change replanning rule to selectively revise task assignments and preserve plan stability even under human interventions. We validate the framework on a 27-component timber-frame assembly. The PSS module achieves 97% state synthesis accuracy, and the HPR module maintains feasible task progression across diverse HRC scenarios. Results indicate that integrating VLM-based perception with knowledge-driven planning improves robustness of state estimation and task planning under dynamic conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。