分离感知与推理,用更小模型高效训练布料操作机器人。
Disentangling perception and reasoning for improving data efficiency in learning cloth manipulation without demonstrations
- 将感知和决策模块解耦,降低模型复杂度。
- 在仿真中训练,仅用小模型即达更好性能。
- 适合资源有限但需高效布料操作的机器人研究者。
布料操作是日常生活中常见但对机器人仍具挑战的任务,其难点在于高维状态空间、复杂动态及易自遮挡。传统解析方法难以提供鲁棒通用策略,强化学习(RL)被视为有前景的解决方案。然而,现有数据驱动方法通常依赖大模型和长时间训练,计算成本高。由于状态估计困难,现有策略多采用端到端学习,以工作区图像为输入,虽支持模拟到现实的迁移,但因环境状态表示损失严重,仍导致高昂计算开销。本文提出一种高效、模块化的RL框架,通过合理设计,显著减少仿真训练中的模型规模与训练时间。实验在SoftGym基准上验证,本方法在任务表现上优于现有基线,且使用更小模型实现。
原文摘要 · Abstract (English)
Cloth manipulation is a ubiquitous task in everyday life, but it remains an open challenge for robotics. The difficulties in developing cloth manipulation policies are attributed to the high-dimensional state space, complex dynamics, and high propensity to self-occlusion exhibited by fabrics. As analytical methods have not been able to provide robust and general manipulation policies, reinforcement learning (RL) is considered a promising approach to these problems. However, to address the large state space and complex dynamics, data-based methods usually rely on large models and long training times. The resulting computational cost significantly hampers the development and adoption of these methods. Additionally, due to the challenge of robust state estimation, garment manipulation policies often adopt an end-to-end learning approach with workspace images as input. While this approach enables a conceptually straightforward sim-to-real transfer via real-world fine-tuning, it also incurs a significant computational cost by training agents on a highly lossy representation of the environment state. This paper questions this common design choice by exploring an efficient and modular approach to RL for cloth manipulation. We show that, through careful design choices, model size and training time can be significantly reduced when learning in simulation. Furthermore, we demonstrate how the resulting simulation-trained model can be transferred to the real world. We evaluate our approach on the SoftGym benchmark and achieve significant performance improvements over available baselines on our task, while using a substantially smaller model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。