arXiv:2509.16611cs.RO2025-09被引 1

从人类示范视频生成可反应的机器人装配行为树

Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly

  • 用视觉语言模型分解示范视频为子任务,生成结构化行为树
  • 实测在长时序装配中成功率超90%,扰动下仍稳定运行
  • 适合需要灵活适应新产品的工业自动化场景

现代制造要求机器人装配系统具备更高灵活性与可靠性。传统方法依赖专家针对每种产品编写固定程序,难以应对产品变更且对环境变化鲁棒性差。由于行为树(BTs)在机器人中兼具模块化与反应性优势,我们提出一种新框架Video-to-BT,将高层认知规划与底层反应控制无缝融合,以行为树作为规划输出和执行控制结构。该方法利用视觉语言模型(VLM)将人类示范视频分解为子任务,并据此生成行为树;执行过程中,计划好的行为树结合实时场景理解,使系统可在动态环境中反应式运行,当执行失败时触发VLM驱动的重规划。闭环架构保障了系统稳定性与适应性。我们在真实装配任务上通过一系列实验验证,结果表明该框架具有高规划可靠性、长时序装配中的稳健表现,以及在多样化扰动条件下的强泛化能力。

原文摘要 · Abstract (English)

Modern manufacturing demands robotic assembly systems with enhanced flexibility and reliability. However, traditional approaches often rely on programming tailored to each product by experts for fixed settings, which are inherently inflexible to product changes and lack the robustness to handle variations. As Behavior Trees (BTs) are increasingly used in robotics for their modularity and reactivity, we propose a novel hierarchical framework, Video-to-BT, that seamlessly integrates high-level cognitive planning with low-level reactive control, with BTs serving both as the structured output of planning and as the governing structure for execution. Our approach leverages a Vision-Language Model (VLM) to decompose human demonstration videos into subtasks, from which Behavior Trees are generated. During the execution, the planned BTs combined with real-time scene interpretation enable the system to operate reactively in the dynamic environment, while VLM-driven replanning is triggered upon execution failure. This closed-loop architecture ensures stability and adaptivity. We validate our framework on real-world assembly tasks through a series of experiments, demonstrating high planning reliability, robust performance in long-horizon assembly tasks, and strong generalization across diverse and perturbed conditions. Project website: https://video2bt.github.io/video2bt_page/

机器人装配行为树视觉语言模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。