arXiv:2509.03219cs.AIcs.LG2025-09中稿 · as a full paper at…被引 2

用不确定性指导探索与利用的切换,提升复杂任务学习效率。

A Novel Framework for Uncertainty-Driven Adaptive Exploration

  • 以不确定性为信号,自动决定何时探索或利用。
  • 在多个环境中表现优于传统自适应探索方法。
  • 兼容多种不确定性测量方式,适用性强。

自适应探索方法通过交替进行探索与利用来学习复杂策略。其关键问题在于如何确定在探索与利用之间切换的合适时机,这在需要学习长而复杂动作序列的任务中尤为关键。本文提出一个通用的自适应探索框架,基于不确定性以合理方式解决该问题。该框架涵盖以往自适应探索方法作为特例,并可集成任意选定的不确定性度量机制,例如内在动机或认知不确定性驱动的探索方法中的机制。实验表明,该框架生成的自适应探索策略在多个环境上均优于标准方法。

原文摘要 · Abstract (English)

Adaptive exploration methods propose ways to learn complex policies via alternating between exploration and exploitation. An important question for such methods is to determine the appropriate moment to switch between exploration and exploitation and vice versa. This is critical in domains that require the learning of long and complex sequences of actions. In this work, we present a generic adaptive exploration framework that employs uncertainty to address this important issue in a principled manner. Our framework includes previous adaptive exploration approaches as special cases. Moreover, we can incorporate in our framework any uncertainty-measuring mechanism of choice, for instance mechanisms used in intrinsic motivation or epistemic uncertainty-based exploration methods. We experimentally demonstrate that our framework gives rise to adaptive exploration strategies that outperform standard ones across several environments.

强化学习自适应探索不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。