提出在线鲁棒强化学习新算法,提升真实场景表现可靠性。
ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning
- 在无先验数据条件下,通过在线交互优化最差情况性能
- 实现子线性后悔率,理论证明接近最优
- 适合模型不准确的真实环境,对机器人等应用友好
针对训练与部署环境分布不匹配导致策略性能下降的问题,本文研究在线分布鲁棒强化学习。在仅有单一未知训练环境、无生成模型或离线数据的条件下,设计了一种基于f-散度不确定性集(如χ²和KL散度球)的高效算法,实现了针对鲁棒控制目标的子线性后悔率,且在最小假设下无需额外数据。同时建立对应极小极大后悔下界,证明方法近似最优。实验表明,在多种存在模型误设的环境中,该方法持续提升最差情况表现,与理论预测一致。
原文摘要 · Abstract (English)
We investigate reinforcement learning (RL) in the presence of distributional mismatch between training and deployment, where policies trained in simulators often underperform in practice due to mismatches between training and deployment conditions, and thereby reliable guarantees on real-world performance are essential. Distributionally robust RL addresses this issue by optimizing worst-case performance over an uncertainty set of environments and providing an optimized lower bound on deployment performance. However, existing studies typically assume access to either a generative model or offline datasets with broad coverage of the deployment environment-assumptions that limit their practicality in unknown environments without prior knowledge. In this work, we study a more practical and challenging setting: online distributionally robust RL, where the agent interacts only with a single unknown training environment while seeking policies that are robust with respect to an uncertainty set around this nominal model. We consider general $f$-divergence-based ambiguity sets, including $χ^2$ and KL divergence balls, and design a computationally efficient algorithm that achieves sublinear regret for the robust control objective under minimal assumptions, without requiring generative or offline data access. Moreover, we establish a corresponding minimax lower bound on the regret of any online algorithm, demonstrating the near-optimality of our method. Experiments across diverse environments with model misspecification show that our approach consistently improves worst-case performance and aligns with the theoretical guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。