arXiv:2603.03146cs.ITcs.AI2026-03

根据信道状态动态调整计算复杂度,提升6G边缘推理吞吐量。

Channel-Adaptive Edge AI: Maximizing Inference Throughput by Adapting Computational Complexity to Channel States

  • 用MvM分布建模高维特征,推导出精度的闭式表达式。
  • 在延迟与精度约束下,使推理吞吐量提升32%以上。
  • 适合6G边缘智能、实时推理场景下的系统设计者。

集成通信与计算(IC²)已成为6G网络中实现高效边缘推理的新范式。然而,由于缺乏可解析的端到端(E2E)推理性能理论框架,相关技术设计面临挑战。该指标需同时考虑信道失真和人工智能模型架构及计算复杂度。本文提出一种可解析的E2E推理精度模型,并据此设计了通道自适应AI算法,以最大化边缘处理速率(EPR),满足延迟与精度约束。考虑服务器部署带早期退出机制的骨干模型,可灵活调整计算复杂度,对移动设备传输的数据特征进行推理。提出的精度模型利用混合冯·米塞斯(Mixture of von Mises, MvM)分布刻画特征在角度域的高维分布,从而获得精度关于量化位宽(反映信道失真)和模型遍历深度(反映计算复杂度)的闭式表达式。基于此模型,求解联合延迟与精度约束下的EPR最大化问题,得到通道自适应算法,实现全链路IC²集成。该算法协同调整发送端特征压缩与接收端模型复杂度,以提升整体效率与推理吞吐量。实验表明,相比固定复杂度方案,其性能显著更优。

原文摘要 · Abstract (English)

\emph{Integrated communication and computation} (IC$^2$) has emerged as a new paradigm for enabling efficient edge inference in sixth-generation (6G) networks. However, the design of IC$^2$ technologies is hindered by the lack of a tractable theoretical framework for characterizing \emph{end-to-end} (E2E) inference performance. The metric is highly complicated as it needs to account for both channel distortion and artificial intelligence (AI) model architecture and computational complexity. In this work, we address this challenge by developing a tractable analytical model for E2E inference accuracy and leveraging it to design a \emph{channel-adaptive AI} algorithm that maximizes inference throughput, referred to as the edge processing rate (EPR), under latency and accuracy constraints. Specifically, we consider an edge inference system in which a server deploys a backbone model with early exit, which enables flexible computational complexity, to perform inference on data features transmitted by a mobile device. The proposed accuracy model characterizes high-dimensional feature distributions in the angular domain using a Mixture of von Mises (MvM) distribution. This leads to a desired closed-form expression for inference accuracy as a function of quantization bit-width and model traversal depth, which represents channel distortion and computational complexity, respectively. Building upon this accuracy model, we formulate and solve the EPR maximization problem under joint latency and accuracy constraints, leading to a channel-adaptive AI algorithm that achieves full IC$^2$ integration. The proposed algorithm jointly adapts transmit-side feature compression and receive-side model complexity according to channel conditions to maximize overall efficiency and inference throughput. Experimental results demonstrate its superior performance as compared with fixed-complexity counterparts.

边缘计算6G自适应推理IC2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。