发现强化学习组件协同效应,提出新框架提升采样效率。
Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control

- 系统分析组件间相互作用,发现堆叠先进方法未必增效。
- 新框架ROSER在连续控制任务中比简单堆叠提升17.60%性能。
- 适合关注高效强化学习系统设计的研究者与工程师。
强化学习系统因内在特性而远比其他机器学习范式复杂,其设计需统筹多个紧密耦合的因素。尽管单个算法组件不断进步,但它们的功能依赖关系仍被忽视:是相互促进还是彼此干扰?我们通过系统性研究发现,不同组件的有效性具有显著的任务依赖性,盲目堆叠前沿技术不仅无法保证性能提升,反而常引发复合非平稳性等新挑战。基于此,我们提炼出可操作的协同设计原则,提出ROSER框架,协调模型基表示、优化稳定性与经验回放三个关键维度。在多种连续控制基准测试中,ROSER持续优于基础模型,相较简单堆叠方案实现17.60%的性能提升。研究强调了强化学习系统设计需全局视角,为开发高样本效率智能体指明方向。
原文摘要 · Abstract (English)
Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many tightly coupled factors. Despite advances in individual algorithmic components, their functional interdependencies remain underexplored: do they exhibit mutual synergy or counterproductive interference? To bridge this gap, we conduct a systematic investigation and find that the efficacy of different components exhibits significant task-dependency, and naively stacking state-of-the-art techniques does not necessarily yield performance gains; instead, it often triggers emergent challenges, such as compounded non-stationarity. Building upon these findings, we distill a suite of actionable insights into the principled coordination of these components. Guided by these insights, we propose ROSER, an RL framework that coordinates three critical dimensions: Model-based Representation, Optimization Stability, and Experience Replay. Across diverse continuous-control benchmarks, ROSER consistently outperforms vanilla baselines and achieves 17.60% gains over naive stack. Our findings underscore the necessity of a holistic perspective in RL system design and paves the way for developing sample-efficient agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。