arXiv:2504.03668cs.DCcs.LG2025-04被引 5

让边缘大模型推理自动适应网络和算力变化,提升响应速度与隐私安全。

Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge

  • 动态拆分模型层并实时迁移,应对边缘环境波动
  • 支持低延迟、高吞吐与隐私保护的多目标优化
  • 适合智慧城市场景、车联网等复杂边缘应用

大型基础模型(LFM)在多模态与生成任务中展现出下一代边缘AI的潜力。然而,在资源受限且异构的边缘环境(如多接入边缘计算,MEC)中执行推理面临巨大挑战,尤其当网络、计算和存储条件随时间变化时。现有分割推理策略无法适应负载波动、带宽动态或隐私约束演变。本文提出一种自适应分割推理编排框架,将模型层的部署位置与分割点设为运行时可调变量。通过三项核心服务实现服务质量(QoS)感知管理:(1) 资源感知的工作负载分发,持续监测节点资源并选择最优的MEC节点子集;(2) 动态分区迁移,透明地重定位已切分的LFM模块以响应利用率或网络变化;(3) 实时重构,动态调整模型层的分割以平衡延迟、吞吐量与隐私。我们形式化了联合部署-分割问题,提出参考架构与算法流程,并讨论其在智慧城市、车联网及工业边缘场景中的适用性。

原文摘要 · Abstract (English)

Large Foundation Models (LFMs), including multi-modal and generative models, promise to unlock new capabilities for next-generation Edge AI applications. However, performing inference with LFMs in resource-constrained and heterogeneous edge environments, such as Multi-access Edge Computing (MEC), presents significant challenges for workload orchestration due to time-varying network, compute, and storage conditions. In particular, current split inference strategies, which partition LFM layers across nodes, are not designed to adapt to fluctuating workloads, dynamic bandwidth conditions, or evolving privacy constraints in high-utilization MEC environments. In this work, we propose a novel adaptive split inference orchestration framework that elevates both the placement and partitioning of LFM layers to runtime-tunable variables. Specifically, our framework enables real-time, quality-of-service (QoS)-aware management of inference workloads by extending conventional orchestrators with three key services: (1) Capacity-aware workload distribution, which continuously profiles node resources and selects an optimal subset of MEC nodes; (2) Dynamic partition migration, which transparently relocates pre-cut LFM segments in response to changes in utilization or network conditions; (3) Real-time reconfiguration, which dynamically re-splits LFM layers to balance latency, throughput, and privacy. We formalize the joint placement-partitioning problem, outline a reference architecture and algorithmic workflow, and discuss applicability in representative smart city, V2X, and industrial edge scenarios.

边缘计算大模型推理动态调度智能编排

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。