arXiv:2412.17344cs.LG2024-12中稿 · AROB-ISBC 2025

让智能体先达标再优化,探索更高效

Reinforcement Learning with a Focus on Adjusting Policies to Reach Targets

  • 以达成目标为优先,动态调整探索强度
  • 在控制与导航任务中表现优于传统方法
  • 适合需要快速适应变化环境的场景

强化学习智能体的目标是通过探索发现更优动作。然而,典型探索策略旨在最大化预期回报,常导致高探索与学习成本。本文提出一种新型深度强化学习方法,优先追求达成特定目标水平,而非最大化期望回报。该方法根据目标达成比例灵活调节探索程度。在运动控制和导航任务上的实验表明,该方法获得的回报不低于或优于标准方法。分析结果显示:本方法能动态调整探索范围,且具备适应非平稳环境的潜力。这些发现表明,该方法在提升强化学习实际应用中的探索效率方面具有有效性。

原文摘要 · Abstract (English)

The objective of a reinforcement learning agent is to discover better actions through exploration. However, typical exploration techniques aim to maximize rewards, often incurring high costs in both exploration and learning processes. We propose a novel deep reinforcement learning method, which prioritizes achieving an aspiration level over maximizing expected return. This method flexibly adjusts the degree of exploration based on the proportion of target achievement. Through experiments on a motion control task and a navigation task, this method achieved returns equal to or greater than other standard methods. The results of the analysis showed two things: our method flexibly adjusts the exploration scope, and it has the potential to enable the agent to adapt to non-stationary environments. These findings indicated that this method may have effectiveness in improving exploration efficiency in practical applications of reinforcement learning.

强化学习探索策略目标导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。