arXiv:2605.09623cs.DCcs.AI2026-05

动态划分神经网络层,实现边缘云环境的节能降时。

Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum

  • 根据运行时状态实时分割模型层,适应硬件与网络变化。
  • 相比静态划分,能耗降低27.09%~35.82%,延迟减少6.34%~22.92%。
  • 适用于资源受限设备上的AI推理,实测验证效果显著。

近年来,人工智能在资源受限的物联网设备上应用日益广泛。然而,现有边缘-云连续体中的AI任务划分与卸载方法多采用静态策略,忽略运行时动态变化,且常在模拟环境中评估而非真实硬件。为此,我们提出一种动态框架,将神经网络层跨异构连续体进行智能分割。该框架在启动时对模型进行剖析,测量节点间网络状况,并周期性重新评估划分以适应环境变化。我们构建了由树莓派边缘设备、笔记本雾节点和高性能桌面云组成的物理测试平台,评估了VGG16、AlexNet和MobileNetV2三个主流卷积神经网络。结果表明,与静态划分基线相比,该框架在能耗上降低27.09%至35.82%,端到端延迟减少6.34%至22.92%,证实了自适应划分优于静态方案。

原文摘要 · Abstract (English)

In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to AI task partitioning and offloading across the edge-cloud continuum typically rely on static methods that ignore runtime dynamics. Furthermore, they are often evaluated in simulated environments rather than on real hardware. To address this gap, we propose a framework that dynamically splits neural network layers across the heterogeneous continuum. The framework profiles the model at startup, measures network link conditions between nodes, and periodically re-evaluates the partition to adapt to environmental changes. We created a physical testbed comprising a Raspberry Pi edge device, a laptop fog, and a high-performance desktop PC as the cloud. We evaluated the framework over three widely adopted convolutional neural networks: VGG16, AlexNet, and MobileNetV2. Our results show that the framework achieves reductions in energy and end-to-end latency of 27.09--35.82% and 6.34--22.92%, respectively, compared to a static partitioning baseline. These findings confirm the superiority of adaptive to static partitioning.

边缘计算AI部署动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。