arXiv:2608.30640cs.LG2026-09

用动作块替代单步动作,显著提升自监督强化学习性能

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

  • 将对比强化学习扩展到动作块级别建模
  • 在18个离线环境上提升31.7%,11个在线环境上提升93.1%
  • 动作块蕴含更多目标信息,改善评估器表征能力

尽管自监督强化学习通过学习状态和动作的表征已取得优异成果,但一个关键开放问题是动作应以何种时间尺度建模。本文脱离传统单步动作范式,将对比强化学习(CRL)扩展至动作块级别,发现这在多个主流离线与在线基准上带来显著且普遍的性能提升:分别在18个和11个环境中实现+31.7%和+93.1%的增益。尽管通常认为动作块有助于建模非马尔可夫、时序扩展的策略并传播无偏多步回报,但实验表明这些解释仅部分适用于CRL。我们的实证研究发现,在CRL背景下,动作块携带的关于目标的信息比单个动作更丰富,从而显著改善了评估器的表征能力,使算法更加有效。

原文摘要 · Abstract (English)

While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open question is the time scale over which actions should be modeled. Departing from the standard formulation relying on single-step actions, we extend contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and find that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively. While action-chunking-driven gains are generally explained through the ability to model non-Markovian, temporally extended policies, and to propagate unbiased multi-step returns, interestingly, we find that these arguments only partially apply to CRL. Our empirical studies suggest that, in the context of CRL, an action chunk carries more information about the goal than a single action, measurably improving the critic's representations, and rendering the algorithm significantly more effective.

强化学习自监督动作建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。