arXiv:2512.01039cs.DCcs.LG2025-12被引 2

动态拆分与部署大模型,让边缘AI实时自适应环境变化。

Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI

  • 运行时动态划分模型层并决定部署位置
  • 在6G边缘环境中实现低延迟与高可用推理
  • 适合需要实时响应的边缘智能系统

在异构边缘环境中对大规模基础模型进行推理,要求一种可重构的编排架构。静态划分模型层依赖计算与网络资源的长期稳定,这与真实部署中的资源波动不匹配。本文提出一个框架,将模型的空间部署与内部切分均作为运行时可解析的变量。编排问题被建模为受延迟、资源利用率和隐私梯度约束的分层分配优化问题。通过融合模型感知的容量分析、动态图重划分与再分配,实现对基础设施波动的响应式推理组合。文中引入了相应的架构与算法组件,并以6G多接入边缘计算场景为例验证可行性。

原文摘要 · Abstract (English)

Inference over large-scale foundation models within heterogeneous edge environments necessitates a fundamentally reconfigurable orchestration substrate. Static partitioning of model layers presumes temporal stability across compute and network resources, which is misaligned with the volatility of real-world deployments. We introduce a framework in which both the spatial placement and internal segmentation of foundation models are elevated to runtime-resolved constructs. The orchestration problem is formalized as a constrained optimization over layer-wise assignments, subject to evolving latency, utilization, and privacy gradients. The framework implements reactive inference composition responsive to infrastructural fluctuations by integrating model-aware capacity profiling with dynamic graph re-partitioning and reallocation. We introduce architectural and algorithmic components, along with a representative use case in 6G multi-access edge computing.

边缘AI模型部署动态调度6G

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。