用合成数据训练视觉系统,让机器人靠褶皱和关键点抓捏折叠衣物。
Synthetic Data Generation and Vision-based Wrinkle and Keypoint Detection for Bimanual Cloth Manipulation

- 用Blender生成带自动标注的合成数据,融合真实图像训练
- 关键点定位误差仅1.76像素,褶皱检测在遮挡下仍有效
- 适合需要双手操作织物的机器人场景,如智能熨烫
机器人操控纺织品仍具挑战,因持续形变和自遮挡影响视觉感知。为解决真实数据标注不足问题,我们开发了基于Blender的合成数据管道,可输出自动标注的关键点,并结合人工标注的渲染图与真实数据训练褶皱检测器。提出感知框架:采用CNN实现排列不变的关键点检测,结合YOLOv8-OpenCV提取结构褶皱上的抓取点。设计双臂算法,先通过褶皱展开完全折叠的衣物,待角部显现后转为关键点引导的熨烫。关键点模型达到1.7615像素的平均位置误差(MPE)。该系统无需微调即可迁移到真实织物,优于基线方法——后者在高遮挡下失效或产生严重折叠误报。
原文摘要 · Abstract (English)
Robotic manipulation of textiles remains challenging because continuous deformation and self-occlusions hinder the robust visual perception required to estimate the cloth's state. To address the lack of annotated real-world data, we developed a Blender-based synthetic pipeline exporting auto-annotated keypoints, and combined manually labeled renders with real-world data to train a wrinkle detector. We present a perception framework integrating a CNN for permutation-invariant keypoint detection and a YOLOv8-OpenCV pipeline to extract grasping points from structural wrinkles. A proposed bimanual algorithm uses this system to stretch fully folded garments via wrinkles, transitioning to keypoint-based ironing once corners emerge. The keypoint model achieves a Mean Position Error (MPE) of 1.7615 pixels. The perception system transfers to physical fabrics without fine-tuning, outperforming baselines that fail in high-occlusion states or yield false positives on severe folds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。