过度模拟真实会阻碍策略学习,新范式可缓解此问题
Too Much of a Good Thing: When sim2real Efforts Impede Policy Learning (And What to Do About It)

- 用sim2sim2real范式,仅以机器人运动学为约束
- 解决仿真过度逼近真实导致的探索受限问题
- 适合关注硬件部署的强化学习研究者
尽管模拟到现实(sim2real)是实现策略有效迁移至硬件的必要手段,但过度追求真实感反而带来负面影响。本文指出,当前的sim2real努力已导致与策略学习目标错位,造成仿真锁定(simulator lock-in)和策略探索能力下降,原因在于现实世界施加了不合理的约束。文章诊断了该问题现状,并提出一种新的sim2sim2real范式:仅将机器人的运动学特性作为设计约束,从而在保持仿真实验灵活性的同时,仍能实现向真实硬件的有效迁移。
原文摘要 · Abstract (English)
While sim2real efforts are necessary for effective policy transfer to hardware, there is such a thing as too much of a good thing. We argue that sim2real efforts have led to misaligned incentives with policy learning, resulting in simulator lock in and poor policy exploration due to the unreasonable constraints imposed by the real world. We offer a diagnosis and explanation of the current status of the problem, and propose a potential solution via a sim2sim2real paradigm that leverages the robot's kinematics as the sole design constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。