arXiv:2505.07841cs.NIcs.LG2025-05被引 4

针对多用户场景下的长序列令牌传输问题,提出高效压缩与优化方案。

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

  • 分两阶段训练:跨模态对齐+任务导向微调,提升通信效率
  • 滑动窗口池化压缩令牌,降低带宽与功耗,延迟下降30%
  • 联合优化资源分配,在不同信噪比下表现更优,适合边缘智能应用

随着基于大模型的智能体兴起,基于Transformer的架构不可避免地产生过长的令牌嵌入,导致高带宽开销、增加功耗和延迟。本文提出一种面向任务的多模态令牌传输方案,以实现高效的多模态信息融合与利用。为提升令牌传输效率,设计了包含跨模态对齐与任务导向微调的两阶段训练算法,并采用滑动窗口池化操作进行令牌压缩,以节省通信资源。为平衡压缩带来的延迟与模型性能之间的权衡,构建了关于延迟与验证损失的加权和优化问题。通过交替优化方法联合优化各用户的带宽、功耗分配与令牌长度。仿真结果表明,所提算法在不同带宽与功耗预算下均优于基线方法;且在多种信噪比条件下,两阶段训练算法的准确率高于无跨模态对齐的方法。

原文摘要 · Abstract (English)

With the emergence of large model-based agents, widely adopted transformer-based architectures inevitably produce excessively long token embeddings for transmission, which may result in high bandwidth overhead, increased power consumption and latency. In this letter, we propose a task-oriented multimodal token transmission scheme for efficient multimodal information fusion and utilization. To improve the efficiency of token transmission, we design a two-stage training algotithm, including cross-modal alignment and task-oriented fine-tuning, for large model-based token communication. Meanwhile, token compression is performed using a sliding window pooling operation to save communication resources. To balance the trade-off between latency and model performance caused by compression, we formulate a weighted-sum optimization problem over latency and validation loss. We jointly optimizes bandwidth, power allocation, and token length across users by using an alternating optimization method. Simulation results demonstrate that the proposed algorithm outperforms the baseline under different bandwidth and power budgets. Moreover, the two-stage training algorithm achieves higher accuracy across various signal-to-noise ratios than the method without cross-modal alignment.

多模态通信令牌压缩边缘计算资源优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。