无需离线调优,实时找到边缘设备的高效能低功耗配置
Covariance-Guided Resource Adaptive Learning for Efficient Edge Inference
- 用距离协方差捕捉硬件参数与性能的非线性关系
- 在单目标场景下达到最优性能的96%~100%
- 适合资源受限设备上的模型部署优化
针对边缘设备上的深度学习推理,相同吞吐量下不同硬件配置的功耗可相差2倍,但用户常难以找到高效配置,且现有方法依赖低效静态设置或需昂贵离线调优。为此,本文提出CORAL,一种无需离线调优的在线优化方法。CORAL利用距离协方差统计捕获硬件参数(如DVFS、并发级别)与性能指标间的非线性依赖关系。不同于以往工作,我们明确定义为吞吐量-功耗协同优化问题,同时满足功耗预算与吞吐量目标。在两个NVIDIA Jetson设备上,对三类从轻量到重型的目标检测模型进行评估。在单目标场景中,CORAL实现96%~100%的最优性能;在严格双约束场景中,当基线失败或超功耗时,CORAL仍能在线持续发现合规配置,探索开销极小。
原文摘要 · Abstract (English)
For deep learning inference on edge devices, hardware configurations achieving the same throughput can differ by 2$\times$ in power consumption, yet operators often struggle to find the efficient ones without exhaustive profiling. Existing approaches often rely on inefficient static presets or require expensive offline profiling that must be repeated for each new model or device. To address this problem, we present CORAL, an online optimization method that discovers near-optimal configurations without offline profiling. CORAL leverages distance covariance to statistically capture the non-linear dependencies between hardware settings, e.g., DVFS and concurrency levels, and performance metrics. Unlike prior work, we explicitly formulate the challenge as a throughput-power co-optimization problem to satisfy power budgets and throughput targets simultaneously. We evaluate CORAL on two NVIDIA Jetson devices across three object detection models ranging from lightweight to heavyweight. In single-target scenarios, CORAL achieves 96% $\unicode{x2013}$ 100% of the optimal performance found by exhaustive search. In strict dual-constraint scenarios where baselines fail or exceed power budgets, CORAL consistently finds proper configurations online with minimal exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。