用视觉域适应让机器人零样本直接上真实场景,成功率超95%
Sim-to-Real Transfer via a Style-Identified Cycle Consistent Generative Adversarial Network: Zero-Shot Deployment on Robotic Manipulators through Visual Domain Adaptation
- 设计SICGAN网络,把虚拟图像转成逼真风格,融合虚实训练
- 虚拟训练成功率达90%-100%,真实部署零样本仍超95%准确率
- 无需调参,适配不同颜色形状物体,适合工业机器人快速落地
深度强化学习(DRL)在工业应用中受限于真实世界训练成本高、样本效率低。虚拟环境虽可降低训练成本,但虚拟与现实间的差异导致策略迁移困难。本文提出基于风格识别循环一致生成对抗网络(SICGAN)的域适应方法,将原始虚拟观测转换为具有真实感的合成图像,构建虚实混合训练环境。经虚拟环境训练后,智能体可直接部署于真实场景,无需额外调优。在两个工业机器人抓取-放置任务中验证:虚拟环境成功率90%至100%,真实部署实现零样本迁移,多数工作区精度超过95%。通过增强现实标记提升评估效率,并实证证明智能体可泛化至不同颜色与形状的实物,包括乐高积木和马克杯。该方案为解决模拟到现实迁移问题提供了高效可扩展的路径。
原文摘要 · Abstract (English)
The sample efficiency challenge in Deep Reinforcement Learning (DRL) compromises its industrial adoption due to the high cost and time demands of real-world training. Virtual environments offer a cost-effective alternative for training DRL agents, but the transfer of learned policies to real setups is hindered by the sim-to-real gap. Achieving zero-shot transfer, where agents perform directly in real environments without additional tuning, is particularly desirable for its efficiency and practical value. This work proposes a novel domain adaptation approach relying on a Style-Identified Cycle Consistent Generative Adversarial Network (StyleID-CycleGAN or SICGAN), an original Cycle Consistent Generative Adversarial Network (CycleGAN) based model. SICGAN translates raw virtual observations into real-synthetic images, creating a hybrid domain for training DRL agents that combines virtual dynamics with real-like visual inputs. Following virtual training, the agent can be directly deployed, bypassing the need for real-world training. The pipeline is validated with two distinct industrial robots in the approaching phase of a pick-and-place operation. In virtual environments agents achieve success rates of 90 to 100\%, and real-world deployment confirms robust zero-shot transfer (i.e., without additional training in the physical environment) with accuracies above 95\% for most workspace regions. We use augmented reality targets to improve the evaluation process efficiency, and experimentally demonstrate that the agent successfully generalizes to real objects of varying colors and shapes, including LEGO\textsuperscript{\textregistered}~cubes and a mug. These results establish the proposed pipeline as an efficient, scalable solution to the sim-to-real problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。