arXiv:2602.13052cs.LGeess.SP2026-02被引 10

为边缘智能体设计量化协同推理,平衡精度、延迟与能耗。

Quantization-Aware Collaborative Inference for Large Embodied AI Models

  • 基于可计算近似建模量化带来的推理失真
  • 推导出比特率与失真之间的上下界关系
  • 联合优化比特数与计算频率,适合资源受限设备

大型人工智能模型(LAIMs)正日益成为具身智能应用的核心智能引擎。然而,其庞大的参数量和计算需求对资源受限的具身智能体构成严峻挑战。为此,本文研究具身智能系统中的量化感知协同推理(co-inference)。首先,提出一种可计算的量化引入推理失真近似方法;基于此,推导出量化比特率-失真函数的上下界,揭示其与模型统计特性(包括量化比特数)的关系。接着,在延迟与能耗约束下,构建联合量化比特数与计算频率的设计问题,旨在最小化失真上界,同时通过对应下界保证紧致性。大量实验验证了所提失真近似、速率-失真边界的准确性,以及联合设计的有效性。仿真与真实测试平台实验表明,该方法在边缘具身智能系统中有效平衡了推理质量、延迟与能耗。

原文摘要 · Abstract (English)

Large artificial intelligence models (LAIMs) are increasingly regarded as a core intelligence engine for embodied AI applications. However, the massive parameter scale and computational demands of LAIMs pose significant challenges for resource-limited embodied agents. To address this issue, we investigate quantization-aware collaborative inference (co-inference) for embodied AI systems. First, we develop a tractable approximation for quantization-induced inference distortion. Based on this approximation, we derive lower and upper bounds on the quantization rate-inference distortion function, characterizing its dependence on LAIM statistics, including the quantization bit-width. Next, we formulate a joint quantization bit-width and computation frequency design problem under delay and energy constraints, aiming to minimize the distortion upper bound while ensuring tightness through the corresponding lower bound. Extensive evaluations validate the proposed distortion approximation, the derived rate-distortion bounds, and the effectiveness of the proposed joint design. Particularly, simulations and real-world testbed experiments demonstrate the effectiveness of the proposed joint design in balancing inference quality, latency, and energy consumption in edge embodied AI systems.

量化协同推理边缘计算具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。