arXiv:2409.18082cs.ROcs.AI2024-09被引 4

用视觉语言模型统一操控各类衣物,提升机器人抓取精度。

SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation

  • 融合视觉与语义信息,用单模型预测不同衣物的关键点轨迹。
  • 在合成数据上训练,关键点检测准确率显著提升,任务成功率更高。
  • 适合希望实现通用衣物操作的机器人研究者与开发者。

自动化衣物操作对辅助机器人构成重大挑战,源于衣物种类多样且易变形。传统方法需为每种衣物单独建模,限制了可扩展性与适应性。本文提出一种统一方法,利用视觉语言模型(VLMs)提升多种衣物类别中的关键点预测性能。通过同时解析视觉与语义信息,该模型使机器人能以单一模型应对不同衣物状态。我们采用先进仿真技术构建大规模合成数据集,实现无需大量真实数据的可扩展训练。实验表明,基于VLM的方法显著提升了关键点检测准确率与任务成功率,为机器人衣物操作提供了更灵活、通用的解决方案。此外,本研究还揭示了VLM在统一多种衣物操作任务方面的潜力,为未来家庭自动化与辅助机器人应用铺平道路。

原文摘要 · Abstract (English)

Automating garment manipulation poses a significant challenge for assistive robotics due to the diverse and deformable nature of garments. Traditional approaches typically require separate models for each garment type, which limits scalability and adaptability. In contrast, this paper presents a unified approach using vision-language models (VLMs) to improve keypoint prediction across various garment categories. By interpreting both visual and semantic information, our model enables robots to manage different garment states with a single model. We created a large-scale synthetic dataset using advanced simulation techniques, allowing scalable training without extensive real-world data. Experimental results indicate that the VLM-based method significantly enhances keypoint detection accuracy and task success rates, providing a more flexible and general solution for robotic garment manipulation. In addition, this research also underscores the potential of VLMs to unify various garment manipulation tasks within a single framework, paving the way for broader applications in home automation and assistive robotics for future.

机器人操作视觉语言模型衣物抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。