arXiv:2604.10974cs.LGcs.RO2026-04被引 2

提出新方法提升强化学习在不确定环境下的鲁棒性与泛化能力。

Robust Adversarial Policy Optimization Under Dynamics Uncertainty

论文配图:Robust Adversarial Policy Optimization Under Dynamics Uncertainty
图 1 · 摘自论文原文
  • 从对偶角度建模鲁棒性与性能权衡,直接优化最坏情况策略。
  • 轨迹级用对抗网络控制温度参数,实现高效稳定最坏情形采样。
  • 模型级通过玻尔兹曼重加权聚焦危险环境,增强政策敏感性覆盖。

强化学习策略在训练分布外的动态环境中常失效,领域随机化及现有对抗强化学习方法未能完全解决此问题。分布鲁棒强化学习虽提供形式化解法,但仍依赖代理对手近似难以求解的原始问题,存在盲区导致不稳定和过度保守。本文提出一种对偶公式,直接揭示鲁棒性与性能的权衡。在轨迹层面,利用对偶问题中的温度参数,通过对抗网络近似,实现受分歧约束下的高效稳定最坏情形滚动;在模型层面,采用动力学集成的玻尔兹曼重加权,聚焦当前策略更不利的环境而非均匀采样。两个组件独立运行且互补:轨迹级调控确保鲁棒滚动,模型级采样提供对策略敏感的恶劣动态覆盖。所提框架鲁棒对抗策略优化(RAPO)优于基准鲁棒强化学习方法,在保持双重可处理性的同时,显著提升对不确定性抗扰能力与分布外动态泛化性能。

原文摘要 · Abstract (English)

Reinforcement learning (RL) policies often fail under dynamics that differ from training, a gap not fully addressed by domain randomization or existing adversarial RL methods. Distributionally robust RL provides a formal remedy but still relies on surrogate adversaries to approximate intractable primal problems, leaving blind spots that potentially cause instability and over-conservatism. We propose a dual formulation that directly exposes the robustness-performance trade-off. At the trajectory level, a temperature parameter from the dual problem is approximated with an adversarial network, yielding efficient and stable worst-case rollouts within a divergence bound. At the model level, we employ Boltzmann reweighting over dynamics ensembles, focusing on more adverse environments to the current policy rather than uniform sampling. The two components act independently and complement each other: trajectory-level steering ensures robust rollouts, while model-level sampling provides policy-sensitive coverage of adverse dynamics. The resulting framework, robust adversarial policy optimization (RAPO) outperforms robust RL baselines, improving resilience to uncertainty and generalization to out-of-distribution dynamics while maintaining dual tractability.

强化学习鲁棒性对抗训练动态不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。