arXiv:2607.09063cs.LG2026-07中稿 · version被引 1

用可自进化预测器,精准估算边缘设备推理延迟。

EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems

论文配图:EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems
图 1 · 摘自论文原文
  • 设计自演化延迟预测框架,边压缩边优化预测精度。
  • 在3个设备、4种模型上优于现有方法,延迟预测更准。
  • 适合需严格延迟约束的实时边缘部署场景。

边缘设备日益用于嵌入式系统中的深度学习应用部署。由于许多应用具有实时性要求且边缘设备资源有限,必须进行针对延迟的神经网络压缩。然而,在真实设备上测量延迟既困难又昂贵。为此,本文提出一种新颖高效的框架EvoLP,可准确预测模型在边缘设备上的推理延迟。该预测器可在网络压缩过程中持续演化,提升预测精度。实验结果表明,EvoLP在三个边缘设备和四种模型变体上均优于现有最先进方法。此外,将其集成到模型压缩框架中,能有效指导压缩过程,在满足严格延迟约束的同时提升模型准确率。代码已开源:https://github.com/ntuliuteam/EvoLP。

原文摘要 · Abstract (English)

Edge devices are increasingly utilized for deploying deep learning applications on embedded systems. The real-time nature of many applications and the limited resources of edge devices necessitate latency-targeted neural network compression. However, measuring latency on real devices is challenging and expensive. Therefore, this letter presents a novel and efficient framework, named EvoLP, to accurately predict the inference latency of models on edge devices. This predictor can evolve to achieve higher latency prediction precision during the network compression process. Experimental results demonstrate that EvoLP outperforms previous state-of-the-art approaches by being evaluated on three edge devices and four model variants. Moreover, when incorporated into a model compression framework, it effectively guides the compression process for higher model accuracy while satisfying strict latency constraints. We open source EvoLP at https://github.com/ntuliuteam/EvoLP.

模型压缩延迟预测边缘计算自演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。