提出生存强化学习,解决长时序规划难题,性能超越对比学习2至8倍。
Survival Reinforcement Learning: Toward Scalable Self-Supervised RL

- 用在线分类替代对比损失,通过延长目标停留时间优化策略。
- 在复杂运动任务中表现优异,长时序任务效果提升2到8倍。
- 适合需要稳定长期规划的机器人系统,为规模化RL提供新思路。
尽管自监督对比强化学习(CRL)展现出显著的深度扩展能力,成功使用超过64层网络,但其在长时程目标条件规划中仍受制于对比损失固有的均匀性容忍困境。本文提出生存强化学习(SRL),一种基于在线分类的替代方法,通过最大化智能体在目标位置的停留时间来扩展生存价值学习框架。SRL规避了CRL的结构限制,并缓解了生存框架中常见的‘开关式’控制问题,该问题常导致复杂动力系统产生不良行为。在多种机器人基准测试中,扩增后的SRL在操作任务上达到与先进CRL相当的性能,在稳定、长时程运动任务中表现更优,性能提升达2至8倍。结果表明,基于分类的方法可能成为推动强化学习规模化的重要基础组件。
原文摘要 · Abstract (English)
While self-supervised Contrastive Reinforcement Learning (CRL) has shown remarkable depth-scaling capabilities, successfully using networks over 64 layers, scaled CRL still struggles with long-horizon goal-conditioned planning due to the uniformity-tolerance dilemma inherent in contrastive losses. We introduce Survival Reinforcement Learning (SRL), an online classification-based alternative that extends the survival value learning framework by maximizing the agent's dwell time at target goals. SRL bypasses the structural constraints of CRL and mitigates the "bang-bang" control solutions inherent to survival frameworks, which often induce undesirable behavior in complex dynamical systems. Evaluated across diverse robotic benchmarks, scaled SRL matches state-of-the-art CRL on manipulation tasks and outperforms it by 2x to 8x on stable, long-horizon locomotion tasks. Our results provide strong additional evidence that classification-based methods may serve as a key primitive in the broader effort to scale reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。