提出分布式多目标强化学习路由,动态适应不同优先级需求。
Dynamic and Distributed Routing in IoT Networks based on Multi-Objective Q-Learning
- 并行学习多种偏好下的Q表,通过贪婪插值快速响应新需求。
- 实测能耗降低80%-90%,交付率提升2-5倍,累积奖励显著更高。
- 适合应急通信、长期监测等需动态调整策略的物联网场景。
物联网网络常面临包交付率、延迟和电池能耗等相互冲突的路由目标。这些优先级可能动态变化:例如紧急告警需要高可靠性,而常规监控则强调节能以延长网络寿命。现有方法多为集中式且假设目标静态,难以快速适应偏好转变。本文提出一种完全分布式的多目标Q-learning路由算法,可并行学习多种偏好对应的Q表,并引入新颖的贪婪插值策略,在无需重训练或中心协调的情况下,对未见偏好实现近优决策。理论分析表明,最优值函数在偏好参数上满足Lipschitz连续性,保证了插值策略的近优性。仿真结果显示,该方法在动态分布式环境下能实时适应优先级变化,在能耗降低80%-90%的同时,累计奖励提升超2-5倍,包交付率显著提高。跨不同偏好窗口长度的敏感性分析也证实,所提DPQ框架始终优于所有基线方法,表现出强鲁棒性。
原文摘要 · Abstract (English)
IoT networks often face conflicting routing goals such as maximizing packet delivery, minimizing delay, and conserving limited battery energy. These priorities can also change dynamically: for example, an emergency alert requires high reliability, while routine monitoring prioritizes energy efficiency to prolong network lifetime. Existing works, including many deep reinforcement learning approaches, are typically centralized and assume static objectives, making them slow to adapt when preferences shift. We propose a dynamic and fully distributed multi-objective Q-learning routing algorithm that learns multiple per-preference Q-tables in parallel and introduces a novel greedy interpolation policy to act near-optimally for unseen preferences without retraining or central coordination. A theoretical analysis further shows that the optimal value function is Lipschitz-continuous in the preference parameter, ensuring that the proposed greedy interpolation policy yields provably near-optimal behavior. Simulations show that our approach adapts in real time to shifting priorities and achieves up to 80-90\% lower energy consumption and more than 2-5x higher cumulative rewards and packet delivery compared to six baseline protocols, under dynamic and distributed settings. Sensitivity analysis across varying preference window lengths confirms that the proposed DPQ framework consistently achieves higher composite reward than all baseline methods, demonstrating robustness to changes in operating conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。