用语言指令指导双臂折叠衣物,实现精准动作生成。
BiFold: Bimanual Cloth Folding with Language Guidance
- 基于视觉语言模型预测文本指令对应的抓取与折叠动作
- 在新衣物和环境中表现优异,成功率达87.3%
- 构建首个带自动标注的双臂折叠数据集,支持语言对齐
衣物折叠因自遮挡、复杂动力学及材质、形状、纹理差异而极具挑战。本文提出BiFold,通过文本指令驱动双臂折叠动作生成,借助预训练视觉-语言模型实现高层抽象指令到具体操作的映射。为解决双臂折叠数据稀缺问题,我们构建了首个自动解析动作并匹配语言指令的新数据集,支持文本条件下的学习。BiFold在现有基准上达到最先进性能,并在新衣物、新指令和新环境间展现出强泛化能力,成功率高达87.3%。
原文摘要 · Abstract (English)
Cloth folding is a complex task due to the inevitable self-occlusions of clothes, their complicated dynamics, and the disparate materials, geometries, and textures that garments can have. In this work, we learn folding actions conditioned on text commands. Translating high-level, abstract instructions into precise robotic actions requires sophisticated language understanding and manipulation capabilities. To do that, we leverage a pre-trained vision-language model and repurpose it to predict manipulation actions. Our model, BiFold, can take context into account and achieves state-of-the-art performance on an existing language-conditioned folding benchmark. To address the lack of annotated bimanual folding data, we introduce a novel dataset with automatically parsed actions and language-aligned instructions, enabling better learning of text-conditioned manipulation. BiFold attains the best performance on our dataset and demonstrates strong generalization to new instructions, garments, and environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。