在上行带宽受限下,用语义多样性提升边缘端传输效果
SAGE: Training-Free Semantic Evidence Composition for Edge-Cloud Inference under Hard Uplink Budgets

- 不依赖训练,结合重要性与嵌入多样性采样
- 传输量减半仍达93%服务器性能上限
- 适合资源受限的边缘计算场景
边缘-云协同推理将复杂输入卸载至远程强大模型,但上行信道对每请求可传输比特数有严格限制。我们发现,仅基于注意力重要性的内容选择策略在硬预算下存在根本局限。两个发现支持此观点:其一,用低重要性但互补的单元替代高重要性单元能提升服务器准确率,说明关键不是单个单元的重要性,而是所传集合对输入多样特征的覆盖程度;其二,在中等预算下,仅按空间均匀采样不依赖内容信息即可获得竞争力准确率,证实空间覆盖本身具有独立价值。基于此分析,我们提出SAGE(语义注意力引导证据组合),一种原则性强、无需训练的方法,融合重要性过滤与嵌入多样性采样。在ImageNet-1K上,SAGE以不足一半可用证据单元的传输量,实现93%的服务器性能上限,显著优于仅依赖重要性的方法。
原文摘要 · Abstract (English)
Edge-cloud hybrid inference offloads difficult inputs to a powerful remote model, but the uplink channel imposes hard per-request constraints on the number of bits that can be transmitted. We show that selecting transmitted content based solely on attention-based importance, the standard approach in collaborative inference, is inherently limited under hard budgets. Two findings support this claim. First, replacing high-importance units with low-importance but complementary ones improves server accuracy. This shows that what matters is not individual importance but how well the transmitted set covers diverse aspects of the input. Second, spatially uniform selection without any content information achieves competitive accuracy at moderate budgets. This confirms that spatial coverage alone carries independent value. Based on this analysis, we propose SAGE (Semantic Attention-Guided Evidence), a principled, training-free method that combines importance filtering with embedding-diversity sampling. SAGE achieves 93% of the server ceiling in offloaded accuracy while transmitting fewer than half of the available evidence units on ImageNet-1K, substantially outperforming importance-only composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。