arXiv:2603.15871cs.LGcs.AI2026-03NeurIPS被引 1

通过反向动作提升强化学习效率,零成本加速训练

Counteractive RL: Rethinking Core Principles for Efficient and Scalable Deep Reinforcement Learning

  • 引入反向动作经验构建新学习范式
  • 在高维环境中实现显著性能提升与样本效率
  • 适合追求高效训练的强化学习研究者

在高维马尔可夫决策过程(MDP)中,强化学习策略面临状态空间指数级增长带来的计算复杂性与策略成功之间的矛盾。本文聚焦学习阶段的智能体-环境交互,提出一种基于反向动作经验的理论驱动新范式。该方法在不增加任何计算开销的前提下,实现了高效、有效、可扩展且加速的学习。我们在包含高维状态表示的雅达利学习环境(ALE)上进行了大量实验,结果验证了理论分析,并在高维环境中显著提升了性能与样本效率。

原文摘要 · Abstract (English)

Following the pivotal success of learning strategies to win at tasks, solely by interacting with an environment without any supervision, agents have gained the ability to make sequential decisions in complex MDPs. Yet, reinforcement learning policies face exponentially growing state spaces in high dimensional MDPs resulting in a dichotomy between computational complexity and policy success. In our paper we focus on the agent's interaction with the environment in a high-dimensional MDP during the learning phase and we introduce a theoretically-founded novel paradigm based on experiences obtained through counteractive actions. Our analysis and method provide a theoretical basis for efficient, effective, scalable and accelerated learning, and further comes with zero additional computational complexity while leading to significant acceleration in training. We conduct extensive experiments in the Arcade Learning Environment with high-dimensional state representation MDPs. The experimental results further verify our theoretical analysis, and our method achieves significant performance increase with substantial sample-efficiency in high-dimensional environments.

强化学习高效训练反向动作样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。