arXiv:2410.10847cs.CVcs.LG2024-10被引 10

用强化学习动态调节边缘设备的算力,让两阶段检测器更稳更快

Lotus: learning-based online thermal and latency variation management for two-stage detectors on edge devices

  • 基于深度强化学习在线联合调节CPU/GPU频率
  • 延迟波动降低超60%,温度下降15℃以上,推理速度提升
  • 适合对实时性要求高的边缘视觉应用

两阶段目标检测器在小物体识别上表现优异,但其高计算开销导致边缘设备热问题加剧,引发运行时频率动态变化,造成显著推理延迟波动。此外,不同帧间候选框数量动态变化进一步加剧计算负载不均,带来更大延迟波动。这种波动严重影响用户体验并浪费硬件资源。为避免热降频并保障稳定推理速度,我们提出Lotus框架,针对两阶段检测器设计基于深度强化学习(DRL)的在线联合频率调控机制,可动态调节CPU和GPU频率。我们在NVIDIA Jetson Orin Nano与小米Mi 11 Lite平台上实现该框架。实验表明,Lotus能持续显著降低延迟波动,在多种场景下实现更快推理速度,并保持更低的CPU/GPU温度。

原文摘要 · Abstract (English)

Two-stage object detectors exhibit high accuracy and precise localization, especially for identifying small objects that are favorable for various edge applications. However, the high computation costs associated with two-stage detection methods cause more severe thermal issues on edge devices, incurring dynamic runtime frequency change and thus large inference latency variations. Furthermore, the dynamic number of proposals in different frames leads to various computations over time, resulting in further latency variations. The significant latency variations of detectors on edge devices can harm user experience and waste hardware resources. To avoid thermal throttling and provide stable inference speed, we propose Lotus, a novel framework that is tailored for two-stage detectors to dynamically scale CPU and GPU frequencies jointly in an online manner based on deep reinforcement learning (DRL). To demonstrate the effectiveness of Lotus, we implement it on NVIDIA Jetson Orin Nano and Mi 11 Lite mobile platforms. The results indicate that Lotus can consistently and significantly reduce latency variation, achieve faster inference, and maintain lower CPU and GPU temperatures under various settings.

边缘计算动态调度强化学习目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。