用离线数据训练调度策略,无需实时交互即可高效处理多用户延迟需求。
Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling
- 基于扩散模型和无采样评判网络,从离线数据学习调度策略
- 在部分可观测与大规模环境中均优于现有方法,性能稳定
- 适合资源敏感、无法在线试错的场景,如数据中心、实时通信
多用户延迟约束调度在具身智能、即时通讯、直播和数据中心管理等实际应用中至关重要,需在用户延迟敏感度差异大、系统动态变化且难以预估的条件下,实现高效资源分配。现有基于学习的方法通常依赖训练阶段与真实系统的在线交互,导致系统性能下降且服务成本高昂。为此,本文提出一种全新的离线强化学习算法SOCD(Scheduling By Offline Learning with Critic Guidance and Diffusion Model),仅利用预先收集的离线数据即可学习高效调度策略。SOCD创新性地结合扩散策略与无采样评判网络,通过将拉格朗日乘子优化融入离线强化学习框架,从可用数据集中高效训练出满足约束的高质量策略,完全避免了对系统在线交互的需求。实验表明,SOCD在部分可观测及大规模环境下的系统动态中均具备强鲁棒性,性能显著优于现有方法。
原文摘要 · Abstract (English)
Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center management, where efficient resource allocation is required among users with diverse delay sensitivities. In these scenarios, schedulers must make real-time decisions to satisfy both delay and resource constraints without prior knowledge of system dynamics, which are often time-varying and challenging to estimate. {Current learning-based methods typically require online interactions with actual systems during the training stage. Therefore, these approaches are often difficult or impractical, as they can significantly degrade system performance and incur substantial service costs.} To address these challenges, we propose a novel offline reinforcement learning-based algorithm, named \underline{S}cheduling By \underline{O}ffline Learning with \underline{C}ritic Guidance and \underline{D}iffusion Model (SOCD), to learn efficient scheduling policies purely from pre-collected \emph{offline data}. SOCD innovatively employs a diffusion policy, complemented by a sampling-free critic network for policy guidance. By integrating the Lagrangian multiplier optimization into the offline reinforcement learning, SOCD efficiently trains high-quality constraint-aware policies exclusively from available datasets, eliminating the need for online interactions with the system. Experimental results demonstrate that SOCD is resilient to various system dynamics, including partially observable and large-scale environments, and delivers superior performance compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。