不训练即可动态合并视觉变压器的令牌,降低计算与传输开销。
Adaptive Pareto-Optimal Token Merging for Edge Transformer Models in Semantic Communication
- 通过多目标优化自动选择每层合并比例,平衡精度与算力。
- 在不同信噪比下显著减少浮点运算量,精度仍具竞争力。
- 支持根据信道质量自适应调整合并强度,灵活权衡延迟与语义保真度。
大规模Transformer模型已成为语义通信系统中提取丰富表征以应对噪声无线信道的强大工具。然而,其巨大的计算需求仍是资源受限的6G网络实际部署的主要障碍。本文提出一种无需训练的自适应令牌合并框架,用于预训练视觉Transformer,旨在同时降低推理时间和传输资源消耗。我们将每层合并比例的选择建模为多目标优化问题,以平衡精度与计算成本。采用基于高斯过程的贝叶斯优化构建帕累托前沿,实现运行时对动态应用需求和信道条件的灵活适应。大量实验表明,该方法持续优于其他基线,在广泛信噪比(SNR)条件下显著减少浮点运算量,同时保持竞争性精度。额外结果表明,根据信道质量自适应调整合并激进程度的策略有效,提供按需权衡延迟与语义保真度的实际机制。这些发现为未来边缘智能系统中部署基于Transformer的语义通信提供了可扩展且高效的解决方案。
原文摘要 · Abstract (English)
Large-scale transformer models have emerged as a powerful tool for semantic communication systems, enabling edge devices to extract rich representations for robust inference across noisy wireless channels. However, their substantial computational demands remain a major barrier to practical deployment in resource-constrained 6G networks. In this paper, we present a training-free framework for adaptive token merging in pretrained vision transformers to jointly reduce inference time and transmission resource usage. We formulate the selection of per-layer merging proportions as a multi-objective optimization problem to balance accuracy and computational cost. We employ Gaussian process-based Bayesian optimization to construct a Pareto frontier of optimal configurations, enabling flexible runtime adaptation to dynamic application requirements and channel conditions. Extensive experiments demonstrate that our method consistently outperforms other baselines and achieves significant reductions in floating-point operations while maintaining competitive accuracy across a wide range of signal-to-noise ratio (SNR) conditions. Additional results highlight the effectiveness of adaptive policies that adjust merging aggressiveness in response to channel quality, providing a practical mechanism to trade off latency and semantic fidelity on demand. These findings establish a scalable and efficient approach for deploying transformer-based semantic communication in future edge intelligence systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。