arXiv:2603.27317cs.RO2026-03

用物理模型梯度引导机器人强化学习探索,更快找到高回报路径。

Where-to-Learn: Analytical Policy Gradient Directed Exploration for On-Policy Robotic Reinforcement Learning

  • 基于可微分动力学模型计算策略梯度,指导探索方向。
  • 在仿真环境中提升策略收敛速度与最终性能,避免无效探索。
  • 适合需要高效探索的机器人控制任务,尤其对物理感知要求高的场景。

在机器人控制中,基于策略的强化学习算法展现了巨大潜力,而有效的探索对于实现高效且高质量的策略学习至关重要。然而,如何高效激励智能体探索更优轨迹仍是一大挑战。现有方法通常通过最大化策略熵或鼓励访问新状态来促进探索,但这些方法忽略了状态潜在价值。本文提出一种新的定向探索机制,利用可微分动力学模型的解析策略梯度,注入任务感知、物理引导的探索指引,从而将智能体导向高回报区域,加速并提升策略学习效果。

原文摘要 · Abstract (English)

On-policy reinforcement learning (RL) algorithms have demonstrated great potential in robotic control, where effective exploration is crucial for efficient and high-quality policy learning. However, how to encourage the agent to explore the better trajectories efficiently remains a challenge. Most existing methods incentivize exploration by maximizing the policy entropy or encouraging novel state visiting regardless of the potential state value. We propose a new form of directed exploration that uses analytical policy gradients from a differentiable dynamics model to inject task-aware, physics-guided guidance, thereby steering the agent towards high-reward regions for accelerated and more effective policy learning.

强化学习机器人控制探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。