用神经科学方法提升边缘计算系统的服务达标率
Benchmarking Dynamic SLO Compliance in Distributed Computing Continuum Systems
- 用主动推理替代传统强化学习动态调节视频参数
- 主动推理在低内存、快收敛上表现更优,延迟达标率超90%
- 适合资源受限的边缘设备实时调度场景
在大规模分布式计算连续体系统(DCCS)中,由于设备异构性和服务需求差异,保障服务等级目标(SLO)面临挑战。不可预测的工作负载与资源限制导致性能波动和SLO违反。本文将新兴的神经科学方法——主动推理,与三种主流强化学习算法(DQN、A2C、PPO)进行对比。以边缘设备运行视频会议与视频流服务器为例,持续监控延迟、带宽等指标,动态调整流数、帧率与分辨率以优化服务质量。通过模拟动态变化的SLO及网络带宽、设备温升等数据漂移场景,结果显示主动推理在内存占用更低、CPU利用稳定、收敛更快方面表现优异,是保障DCCS中SLO合规性的有前景方案。
原文摘要 · Abstract (English)
Ensuring Service Level Objectives (SLOs) in large-scale architectures, such as Distributed Computing Continuum Systems (DCCS), is challenging due to their heterogeneous nature and varying service requirements across different devices and applications. Additionally, unpredictable workloads and resource limitations lead to fluctuating performance and violated SLOs. To improve SLO compliance in DCCS, one possibility is to apply machine learning; however, the design choices are often left to the developer. To that extent, we provide a benchmark of Active Inference -- an emerging method from neuroscience -- against three established reinforcement learning algorithms (Deep Q-Network, Advantage Actor-Critic, and Proximal Policy Optimization). We consider a realistic DCCS use case: an edge device running a video conferencing application alongside a WebSocket server streaming videos. Using one of the respective algorithms, we continuously monitor key performance metrics, such as latency and bandwidth usage, to dynamically adjust parameters -- including the number of streams, frame rate, and resolution -- to optimize service quality and user experience. To test algorithms' adaptability to constant system changes, we simulate dynamically changing SLOs and both instant and gradual data-shift scenarios, such as network bandwidth limitations and fluctuating device thermal states. Although the evaluated algorithms all showed advantages and limitations, our findings demonstrate that Active Inference is a promising approach for ensuring SLO compliance in DCCS, offering lower memory usage, stable CPU utilization, and fast convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。