arXiv:2605.01516cs.RO2026-05

用高保真模拟器训练的动态模型,让强化学习更高效且可迁移。

Dynamics Distillation for Efficient and Transferable Control Learning

论文配图:Dynamics Distillation for Efficient and Transferable Control Learning
图 1 · 摘自论文原文
  • 将高精度仿真器的动态特性提炼成可并行学习的模型
  • 在简化环境中训练的策略能可靠迁移到真实仿真器
  • 评估动态模型需看其生成策略的质量,不只看预测精度

自主驾驶的鲁棒控制策略学习需要训练环境兼具物理真实性和计算可扩展性,而现有模拟器仅能单独满足其一。我们提出 Sim2Sim2Sim 框架,通过将高保真车辆模拟器的动力学特性蒸馏到一个高度可并行化的学习动力学模型中,实现从高保真模拟到可扩展强化学习的桥梁。在该蒸馏环境中纯化训练控制策略,并部署回原始高保真源模拟器,验证了更高效的策略优化和在复杂动态下的可靠迁移能力。进一步表明,仅靠预测精度无法全面衡量学习动力学模型作为强化学习训练环境的适用性,还需评估其生成策略的质量。

原文摘要 · Abstract (English)

Robust control policy learning for autonomous driving requires training environments to be both physically realistic and computationally scalable, properties that existing simulators provide only in isolation. We introduce Sim2Sim2Sim, a framework that bridges high-fidelity vehicle simulation and scalable reinforcement learning by distilling simulator dynamics into a highly parallelizable learned dynamics model. By training control policies purely within this distilled environment and deploying them back into the high-fidelity source simulator, we demonstrate more efficient policy optimization and reliable transfer under challenging dynamics. We further show that predictive accuracy alone does not fully characterize a learned dynamics model's suitability as a reinforcement learning training environment, which should also be assessed by the quality of the policies it enables.

强化学习仿真迁移动态蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。