arXiv:2509.02324cs.RO2025-09被引 5

用大模型规划+视觉语言对齐,让机器人听懂指令叠布料

Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception

  • 用大语言模型拆解指令为可执行动作,打通语义到操作的鸿沟
  • 仿真中对新指令、新任务成功率分别提升33.3%和2.23倍
  • 适合做复杂长序列操作的机器人研究者或工程落地团队

语言引导的长时序可变形物体操作因自由度高、动力学复杂且需精准视觉-语言对齐而极具挑战。本文聚焦多步布料折叠这一典型任务,提出一个统一框架,融合大语言模型(LLM)规划器、视觉-语言模型(VLM)感知系统与任务执行模块。其中,LLM规划器将高层语言指令分解为底层动作原语,弥合语义与执行的差距,实现感知与动作对齐,并增强泛化能力;VLM感知模块采用SigLIP2架构,结合双向交叉注意力融合机制与权重解耦低秩适配(DoRA)微调,实现语言条件下的细粒度视觉定位。仿真与真实场景实验表明,该方法在仿真中对已见指令、未见指令及未见任务的性能分别优于当前最优基线2.23、1.87和33.3个百分点;在真实机器人上,能稳定执行多种材质和构型下的多步折叠序列,展现强大实际泛化能力。

原文摘要 · Abstract (English)

Language-guided long-horizon manipulation of deformable objects presents significant challenges due to high degrees of freedom, complex dynamics, and the need for accurate vision-language grounding. In this work, we focus on multi-step cloth folding, a representative deformable-object manipulation task that requires both structured long-horizon planning and fine-grained visual perception. To this end, we propose a unified framework that integrates a Large Language Model (LLM)-based planner, a Vision-Language Model (VLM)-based perception system, and a task execution module. Specifically, the LLM-based planner decomposes high-level language instructions into low-level action primitives, bridging the semantic-execution gap, aligning perception with action, and enhancing generalization. The VLM-based perception module employs a SigLIP2-driven architecture with a bidirectional cross-attention fusion mechanism and weight-decomposed low-rank adaptation (DoRA) fine-tuning to achieve language-conditioned fine-grained visual grounding. Experiments in both simulation and real-world settings demonstrate the method's effectiveness. In simulation, it outperforms state-of-the-art baselines by 2.23, 1.87, and 33.3 on seen instructions, unseen instructions, and unseen tasks, respectively. On a real robot, it robustly executes multi-step folding sequences from language instructions across diverse cloth materials and configurations, demonstrating strong generalization in practical scenarios. Project page: https://language-guided.netlify.app/

机器人操作大模型规划视觉语言对齐布料折叠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。