arXiv:2602.23722cs.NIcs.AI2026-02中稿 · IEEE INFOCOM Works…被引 1

研究如何在5G网络中跨设备、边缘和云实现毫秒级大模型推理。

SLA-Aware Distributed LLM Inference Across Device-RAN-Cloud

  • 通过分层部署与量化模型优化,提升RAN边缘推理速度。
  • 量化模型可稳定在0.5秒内完成,未量化则常超时。
  • 多实例GPU隔离保障基带处理不被干扰,适合真实场景部署。

具身智能需要在无线接入网(RAN)附近实现亚秒级推理,但实际部署涉及异构层级(设备端、RAN边缘、云端),且不能干扰实时基带处理。我们在一个5G独立组网(SA)的AI-RAN测试平台上进行测量,采用固定基线策略以确保可重复性。系统包含设备端、三节点共宿容器化5G RAN的RAN边缘集群以及云端。结果表明,设备端执行仍需数秒,无法满足亚秒要求;在RAN边缘,服务等级协议(SLA)可行性主要取决于模型选择:量化模型可保持在0.5秒内完成,而未量化及部分较大量化模型因阻塞与排队导致任务超时。在云端,受测广域网路径下,仅32.9%请求能在0.5秒内完成,但所有模型均能在1.0秒内完成(100%达标)。在下行链路饱和且最多20个并发推理客户端条件下,多实例GPU(MIG)隔离能维持基带时序健康指标,支持固定分区下的安全共置。

原文摘要 · Abstract (English)

Embodied AI requires sub-second inference near the Radio Access Network (RAN), but deployments span heterogeneous tiers (on-device, RAN-edge, cloud) and must not disrupt real-time baseband processing. We report measurements from a 5G Standalone (SA) AI-RAN testbed using a fixed baseline policy for repeatability. The setup includes an on-device tier, a three-node RAN-edge cluster co-hosting a containerized 5G RAN, and a cloud tier. We find that on-device execution remains multi-second and fails to meet sub-second budgets. At the RAN edge, SLA feasibility is primarily determined by model variant choice: quantized models concentrate below 0.5\,s, while unquantized and some larger quantized models incur deadline misses due to stalls and queuing. In the cloud tier, meeting a 0.5\,s deadline is challenging on the measured WAN path (up to 32.9\% of requests complete within 0.5\,s), but all evaluated variants meet a 1.0\,s deadline (100\% within 1.0\,s). Under saturated downlink traffic and up to $N{=}20$ concurrent inference clients, Multi-Instance GPU (MIG) isolation preserves baseband timing-health proxies, supporting safe co-location under fixed partitioning.

大模型推理5G AI边缘计算SLA保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。