arXiv:2604.15289cs.RO2026-04被引 1

用抽象仿真器训练机器人,通过历史信息提升真实世界迁移效果。

Abstract Sim2Real through Approximate Information States

论文配图:Abstract Sim2Real through Approximate Information States
图 1 · 摘自论文原文
  • 基于状态抽象理论,用历史状态修正抽象仿真器动态
  • 在仿真到仿真、仿真到真实场景中均实现成功策略迁移
  • 适合仿真不完整但需跨域部署的机器人任务

近年来,强化学习(RL)在机器人领域取得显著进展,前提是存在快速准确的模拟器。然而,随着机器人应用范围扩展至更复杂的大规模场景,完全真实的模拟器越来越难以构建。在此背景下,模拟器往往无法涵盖目标任务的关键细节,这促使我们研究在抽象模拟器上进行模拟到现实的迁移问题。本文正式定义并研究了抽象模拟到现实(Abstract Sim2Real)问题:给定一个在粗粒度层次上建模目标任务的抽象模拟器,如何在该抽象模拟器中使用强化学习训练策略,并成功迁移到真实世界?我们的第一个贡献是利用强化学习中的状态抽象语言形式化该问题。分析表明,若抽象模拟器的动态能纳入状态历史信息,则可实现与目标任务的对齐。基于此理论框架,我们提出一种方法,利用真实任务数据修正抽象模拟器的动态。实验结果表明,该方法在仿真到仿真及仿真到真实场景中均实现了成功的策略迁移。

原文摘要 · Abstract (English)

In recent years, reinforcement learning (RL) has shown remarkable success in robotics when a fast and accurate simulator is available for a given task. When using RL and simulation, more simulator realism is generally beneficial but becomes harder to obtain as robots are deployed in increasingly complex and widescale domains. In such settings, simulators will likely fail to model all relevant details of a given target task and this observation motivates the study of sim2real with simulators that leave out key task details. In this paper, we formalize and study the abstract sim2real problem: given an abstract simulator that models a target task at a coarse level of abstraction, how can we train a policy with RL in the abstract simulator and successfully transfer it to the real-world? Our first contribution is to formalize this problem using the language of state abstraction from the RL literature. This framing shows that an abstract simulator can be grounded to match the target task if the grounded abstract dynamics take the history of states into account. Based on the formalism, we then introduce a method that uses real-world task data to correct the dynamics of the abstract simulator. We then show that this method enables successful policy transfer both in sim2sim and sim2real evaluation.

强化学习模拟到现实状态抽象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。