动态链路下自适应压缩,让远程推理更快更准
Remote Inference over Dynamic Links via Adaptive Rate Deep Task-Oriented Vector Quantization
- 用嵌套码本+渐进学习实现可变速率压缩
- 支持多比特预算,交换更多比特后精度持续提升
- 适合实时远程推理场景,尤其在网络波动时表现优
众多技术依赖远程推理,即本地采集数据通过通信信道传至远端服务器进行推断。由于通信通道常受限于带宽,需对数据进行压缩以降低延迟。深度学习虽能联合设计压缩映射与编码推理规则,但现有学习型压缩方法均为静态,难以适应信道条件变化和动态链路。为此,我们提出自适应速率任务导向向量量化(ARTOVeQ),一种专为动态链路上的远程推理设计的可学习压缩机制。ARTOVeQ基于嵌套码本设计,并采用渐进学习算法。实验表明,该方法支持低延迟推理的逐步精炼,可在传输高维数据时同时使用多种分辨率。数值结果证明,所提方案实现了多速率远程深度推理,在广泛比特预算下运行良好,且随着传输比特数增加,推理精度逐步提升,性能接近单速率深度量化方法。
原文摘要 · Abstract (English)
A broad range of technologies rely on remote inference, wherein data acquired is conveyed over a communication channel for inference in a remote server. Communication between the participating entities is often carried out over rate-limited channels, necessitating data compression for reducing latency. While deep learning facilitates joint design of the compression mapping along with encoding and inference rules, existing learned compression mechanisms are static, and struggle in adapting their resolution to changes in channel conditions and to dynamic links. To address this, we propose Adaptive Rate Task-Oriented Vector Quantization (ARTOVeQ), a learned compression mechanism that is tailored for remote inference over dynamic links. ARTOVeQ is based on designing nested codebooks along with a learning algorithm employing progressive learning. We show that ARTOVeQ extends to support low-latency inference that is gradually refined via successive refinement principles, and that it enables the simultaneous usage of multiple resolutions when conveying high-dimensional data. Numerical results demonstrate that the proposed scheme yields remote deep inference that operates with multiple rates, supports a broad range of bit budgets, and facilitates rapid inference that gradually improves with more bits exchanged, while approaching the performance of single-rate deep quantization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。