arXiv:2506.10133cs.LGcs.RO2025-06

用真实数据优化模拟器参数,提升强化学习的现实部署效果

Statistical Guarantees for Offline Domain Randomization

  • 基于真实数据拟合模拟器参数分布,改进传统域随机化
  • 理论证明参数估计随数据量增长趋于真实值
  • 为离线强化学习提供可信赖的参数随机化依据

强化学习代理在从仿真部署到现实世界时常表现不佳。主流方法域随机化(DR)通过在多个动态参数采样生成的模拟器上训练策略来缩小仿真与现实的差距,但标准DR忽略了已有的真实系统离线数据。本文研究离线域随机化(ODR),先利用离线数据拟合模拟器参数分布。尽管已有大量实证工作报告了如DROPO等算法的显著收益,但其理论基础仍不清晰。本文将ODR建模为参数化模拟器族上的最大似然估计,并给出统计保证:在温和正则性和可识别性条件下,估计器弱收敛(随数据量增加概率收敛至真实动态);若额外满足统一Lipschitz连续性假设,则强收敛(几乎必然收敛)。我们检验了这些假设的实际可行性,并提出放宽条件,支持ODR在更广泛场景中的应用。结果为ODR提供了严谨理论基础,明确了何时可凭离线数据可靠指导下游离线强化学习的随机化分布选择。

原文摘要 · Abstract (English)

Reinforcement-learning (RL) agents often struggle when deployed from simulation to the real-world. A dominant strategy for reducing the sim-to-real gap is domain randomization (DR) which trains the policy across many simulators produced by sampling dynamics parameters, but standard DR ignores offline data already available from the real system. We study offline domain randomization (ODR), which first fits a distribution over simulator parameters to an offline dataset. While a growing body of empirical work reports substantial gains with algorithms such as DROPO, the theoretical foundations of ODR remain largely unexplored. In this work, we cast ODR as a maximum-likelihood estimation over a parametric simulator family and provide statistical guarantees: under mild regularity and identifiability conditions, the estimator is weakly consistent (it converges in probability to the true dynamics as data grows), and it becomes strongly consistent (i.e., it converges almost surely to the true dynamics) when an additional uniform Lipschitz continuity assumption holds. We examine the practicality of these assumptions and outline relaxations that justify ODR's applicability across a broader range of settings. Taken together, our results place ODR on a principled footing and clarify when offline data can soundly guide the choice of a randomization distribution for downstream offline RL.

强化学习域随机化离线学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。