让网络管理智能体动态适应不同优先级,提升多目标适应能力。
Dynamic Preference Multi-Objective Reinforcement Learning for Internet Network Management

- 基于状态和可变偏好选择动作,实现动态决策
- 在多种偏好组合下泛化性能显著优于静态偏好方法
- 提出偏好分布估计方法,支持无偏训练,适合复杂网络场景
互联网服务提供商需在高服务质量(QoS)与最小计算资源使用之间平衡多个目标。传统强化学习方法通常采用固定重要性权重的单一奖励函数,即静态偏好。然而实际中偏好会随网络状态、外部因素变化——例如服务器宕机可能导致连锁过载时,应降低QoS权重、提高资源节约权重。本文提出新型基于强化学习的网络管理智能体,能根据当前状态与动态偏好选择动作,使单个智能体适应多种偏好场景。同时,提出一种数值方法用于估计对训练有利的偏好分布,确保训练无偏。实验表明,该方法在多种偏好配置下的泛化性能显著优于传统静态偏好方法;多个分析也验证了所提偏好估计方法的优势。
原文摘要 · Abstract (English)
An internet network service provider manages its network with multiple objectives, such as high quality of service (QoS) and minimum computing resource usage. To achieve these objectives, a reinforcement learning-based (RL) algorithm has been proposed to train its network management agent. Usually, their algorithms optimize their agents with respect to a single static reward formulation consisting of multiple objectives with fixed importance factors, which we call preferences. However, in practice, the preference could vary according to network status, external concerns and so on. For example, when a server shuts down and it can cause other servers' traffic overloads leading to additional shutdowns, it is plausible to reduce the preference of QoS while increasing the preference of minimum computing resource usages. In this paper, we propose new RL-based network management agents that can select actions based on both states and preferences. With our proposed approach, we expect a single agent to generalize on various states and preferences. Furthermore, we propose a numerical method that can estimate the distribution of preference that is advantageous for unbiased training. Our experiment results show that the RL agents trained based on our proposed approach significantly generalize better with various preferences than the previous RL approaches, which assume static preference during training. Moreover, we demonstrate several analyses that show the advantages of our numerical estimation method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。