用视觉语言模型和强化学习实现机器人自动穿绳,成功率92%。
Hierarchical DLO Routing with Reinforcement Learning and In-Context Vision-language Models
- 分层架构:用VLM进行上下文推理生成多步计划
- 92%成功率,在5段长序列任务中表现稳定
- 支持隐含指令和复杂场景,适合工业装配应用
可变形线性物体(如电缆、绳索)的长时序穿引任务在工业装配和日常生活中常见。这类任务极具挑战性,要求机器人具备长周期规划能力与可靠技能执行能力。成功完成需适应非线性动力学,分解抽象路由目标,并生成由多个技能组成的多步计划,全过程依赖精确的高层推理。本文提出一种完全自主的分层框架,针对语言表达的显式或隐式路由目标,利用视觉语言模型(VLMs)进行上下文高阶推理以合成可行计划,再由强化学习训练的低层技能执行。为提升长时序鲁棒性,引入故障恢复机制,将物体重新定向至可插入状态。方法在包含物体属性、空间描述、隐含语言指令及扩展5片段设置的多样化场景中均表现良好,长时序路由任务整体成功率高达92%。
原文摘要 · Abstract (English)
Long-horizon routing tasks of deformable linear objects (DLOs), such as cables and ropes, are common in industrial assembly lines and everyday life. These tasks are particularly challenging because they require robots to manipulate DLO with long-horizon planning and reliable skill execution. Successfully completing such tasks demands adapting to their nonlinear dynamics, decomposing abstract routing goals, and generating multi-step plans composed of multiple skills, all of which require accurate high-level reasoning during execution. In this paper, we propose a fully autonomous hierarchical framework for solving challenging DLO routing tasks. Given an implicit or explicit routing goal expressed in language, our framework leverages vision-language models~(VLMs) for in-context high-level reasoning to synthesize feasible plans, which are then executed by low-level skills trained via reinforcement learning. To improve robustness over long horizons, we further introduce a failure recovery mechanism that reorients the DLO into insertion-feasible states. Our approach generalizes to diverse scenes involving object attributes, spatial descriptions, implicit language commands, and \myred{extended 5-clip settings}. It achieves an overall success rate of 92\% across long-horizon routing scenarios. Please refer to our project page: https://icra2026-dloroute.github.io/DLORoute/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。