改进网络架构可显著提升多目标强化学习的解集质量。
Momba: Network Modernization Improves Multi-Objective Reinforcement Learning
- 引入归一化与分布回报建模,增强函数逼近能力
- 在连续控制基准上解集质量明显提升,样本效率更高
- 无需改动算法框架,适合追求性能优化的研究者
深度强化学习近期进展表明,优化神经网络结构可在不改变算法的前提下显著提升样本效率和最终性能。然而,多目标强化学习(MORL)主要聚焦于算法创新,对网络架构的关注较少。尽管最优策略和价值函数会随目标权衡变化显著不同,现有MORL算法仍普遍使用简单前馈网络,以权衡参数为条件进行表示。这引发疑问:更强大的函数逼近器能否提升算法表现?本文整合了近期神经网络设计进展:(i) 观测与特征归一化,(ii) 权重归一化,以及 (iii) 基于熵正则化的分布回报建模。在标准连续控制基准上的实验证明,这些改进显著提升了生成解集的质量,且无需对底层算法做重大修改。
原文摘要 · Abstract (English)
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。