arXiv:2507.06111cs.LGcs.RO2025-07被引 3

安全训练强化学习策略,避免真实环境试错

Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation

  • 用多评价值网络量化策略不确定性,识别分布外状态
  • 在仿真中迭代优化高不确定区域,提升真实场景泛化能力
  • 适合需要安全部署的机器人控制、自动驾驶等应用

将强化学习策略部署到真实世界面临分布偏移、安全风险及难以直接交互等问题。现有方法如域随机化(DR)和离动态强化学习需在目标域直接交互,存在安全隐患。本文提出不确定性感知强化学习(UARL),通过集成评价值网络量化策略不确定性,并结合渐进式环境随机化,在不接触真实环境的前提下实现安全训练。通过在仿真环境中对状态空间的高不确定性区域进行迭代优化,提升策略对目标域的鲁棒泛化能力。在MuJoCo基准测试与四足机器人上的实验表明,UARL能可靠检测分布外状态,性能优于基线方法,且样本效率更高。

原文摘要 · Abstract (English)

Deploying reinforcement learning (RL) policies in real-world involves significant challenges, including distribution shifts, safety concerns, and the impracticality of direct interactions during policy refinement. Existing methods, such as domain randomization (DR) and off-dynamics RL, enhance policy robustness by direct interaction with the target domain, an inherently unsafe practice. We propose Uncertainty-Aware RL (UARL), a novel framework that prioritizes safety during training by addressing Out-Of-Distribution (OOD) detection and policy adaptation without requiring direct interactions in target domain. UARL employs an ensemble of critics to quantify policy uncertainty and incorporates progressive environmental randomization to prepare the policy for diverse real-world conditions. By iteratively refining over high-uncertainty regions of the state space in simulated environments, UARL enhances robust generalization to the target domain without explicitly training on it. We evaluate UARL on MuJoCo benchmarks and a quadrupedal robot, demonstrating its effectiveness in reliable OOD detection, improved performance, and enhanced sample efficiency compared to baselines.

强化学习安全训练分布外检测仿真迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。