arXiv:2510.09942cs.LGcs.AI2025-10被引 5

通过稀疏化压缩提升边缘云推理速度,降低带宽压力。

Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding

  • 基于分布稀疏性与共形预测,动态压缩候选词
  • 在不同场景下将延迟降低30%以上,拒稿率下降40%
  • 适合资源受限设备部署,尤其适用于边缘计算

边缘-云推测解码(SD)通过边缘端小型语言模型(SLM)生成草稿标记,由云端大型语言模型(LLM)进行验证,从而加速推理。其核心瓶颈在于边缘-云链路带宽有限,需高效压缩草稿标记分布。本文首先推导出信息论边界,将标记拒绝率分解为SLM-LLM分布不匹配与量化失真两部分贡献。基于此分析,提出稀疏量化与采样框架SQS-SD,利用结构化稀疏化和基于格的量化。其中K-SQS采用固定Top-K截断,而C-SQS则通过在线共形预测自适应调整保留标记集,确保与密集分布偏差有界。实验表明,两种方法在互补工作区间内均有效降低端到端延迟与拒绝率。

原文摘要 · Abstract (English)

Edge-cloud speculative decoding (SD) accelerates inference by having a cloud-based large language model (LLM) that verifies draft tokens generated by a resource-constrained small language model (SLM) at the edge. A central bottleneck is the limited bandwidth of the edge-cloud link, which necessitates efficient compression of draft token distributions. We first derive an information-theoretic bound that decomposes the token rejection rate into contributions from SLM-LLM distribution mismatch and from quantization distortion. Guided by this analysis, we propose the Sparse Quantize-and-Sample SD (SQS-SD) framework, which exploits distributional sparsity through structured sparsification and lattice-based quantization. Within this framework, K-SQS applies fixed top-K truncation, while C-SQS adaptively adjusts the retained token set via online conformal prediction to ensure bounded deviation from the dense distribution. Empirical results confirm that both approaches improve end-to-end latency and rejection rates in complimentary operating regimes.

边缘计算推理加速稀疏化共形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。