arXiv:2503.08372cs.RO2025-03被引 10

用语言指令指导,让机器人自动折叠各类衣物

MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model

  • 分拆任务规划与动作预测,独立学习提升泛化能力
  • 支持多品类衣物折叠,可响应不同语言指令
  • 基于点云轨迹生成与基础模型,适合复杂变形物体操作

衣物折叠是机器人操作中的常见但极具挑战的任务。由于衣物的可变形性,其状态空间庞大且动态复杂,难以实现精确精细的操作。以往方法常依赖预设关键点或示范数据,限制了在多种衣物类别间的泛化能力。本文提出MetaFold框架,将任务规划与动作预测解耦,分别独立学习以增强模型泛化性。该框架采用语言引导的点云轨迹生成进行任务规划,并使用低层基础模型进行动作预测。这种结构支持多类别学习,使模型能灵活适应用户的不同指令和折叠任务。实验结果表明,所提框架具有显著优势。补充材料见:https://meta-fold.github.io/

原文摘要 · Abstract (English)

Garment folding is a common yet challenging task in robotic manipulation. The deformability of garments leads to a vast state space and complex dynamics, which complicates precise and fine-grained manipulation. Previous approaches often rely on predefined key points or demonstrations, limiting their generalization across diverse garment categories. This paper presents a framework, MetaFold, that disentangles task planning from action prediction, learning each independently to enhance model generalization. It employs language-guided point cloud trajectory generation for task planning and a low-level foundation model for action prediction. This structure facilitates multi-category learning, enabling the model to adapt flexibly to various user instructions and folding tasks. Experimental results demonstrate the superiority of our proposed framework. Supplementary materials are available on our website: https://meta-fold.github.io/.

机器人操作语言引导衣物折叠轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。