arXiv:2411.00859cs.LGcs.AI2024-11被引 7

通过分析模型与硬件特性,实现边缘AI计算卸载的高效资源分配。

Profiling AI Models: Towards Efficient Computation Offloading in Heterogeneous Edge AI Systems

  • 基于模型类型、超参数和硬件信息构建性能预测模型。
  • 3000+次实验验证可精准预估资源消耗与任务完成时间。
  • 适合研究边缘计算、6G网络中异构设备协同的开发者。

端侧AI应用(如计算机视觉、生成式AI)的快速发展带来了海量数据与算力需求,常超出终端设备处理能力。边缘AI通过将计算任务卸载至网络边缘来缓解此问题,对6G未来服务至关重要。然而,现有方法面临并发卸载时资源受限、假设系统同质化等挑战。为此,本文提出一项研究路线,聚焦于对AI模型进行性能剖析,采集模型类型、超参数及底层硬件信息,以预测资源占用与任务完成时间。初步实验涵盖3000多次运行,结果表明该方法能有效优化资源分配,提升边缘AI性能。

原文摘要 · Abstract (English)

The rapid growth of end-user AI applications, such as computer vision and generative AI, has led to immense data and processing demands often exceeding user devices' capabilities. Edge AI addresses this by offloading computation to the network edge, crucial for future services in 6G networks. However, it faces challenges such as limited resources during simultaneous offloads and the unrealistic assumption of homogeneous system architecture. To address these, we propose a research roadmap focused on profiling AI models, capturing data about model types, hyperparameters, and underlying hardware to predict resource utilisation and task completion time. Initial experiments with over 3,000 runs show promise in optimising resource allocation and enhancing Edge AI performance.

边缘计算资源调度6G

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。