arXiv:2606.14808eess.IVcs.CV2026-06

让图像通信更懂任务,还能解释决策过程

Explainable Task-Oriented Token Communication for AI-Native 6G Networks

论文配图:Explainable Task-Oriented Token Communication for AI-Native 6G Networks
图 1 · 摘自论文原文
  • 用视觉令牌和任务令牌协同传输,提升任务相关性
  • 通过跨模态注意力机制,实现任务对视觉信息的精准引导
  • 生成注意力热图,可解释任务决策依据,适合可信AI研究者

基础模型与无线通信的融合正推动图像通信从比特级精确传输向任务导向传输演进。然而,现有任务导向图像通信方法仍面临三大挑战:任务导向令牌表征不足、视觉令牌与任务令牌协作不充分、任务决策可解释性差。为此,本文提出可解释的任务导向令牌通信(ET-TokenCom)框架。该框架将令牌视为统一的信息表示与传输单元,构建覆盖视觉感知、无线传输与任务推理的端到端通信链路。在发送端,从图像中提取视觉令牌以保留低层视觉信息,同时引入由基础模型生成的任务令牌,表示当前任务所需的目标信息与决策意图。设计跨模态注意力(CMA)融合机制,使任务令牌能显式引导视觉令牌的选择、加权与传输。在接收端,将令牌解码与可解释输出机制结合,生成注意力热图,突出不同任务目标下的关键感知区域,并揭示任务令牌对输出的影响。仿真结果验证了所提框架的有效性与鲁棒性。

原文摘要 · Abstract (English)

The integration of Foundation Models (FMs) and wireless communications is driving the evolution of image communication from bit-accurate transmission toward task-oriented transmission. However, existing task-oriented image communication methods still face three major challenges: insufficient task-oriented Token representation, inadequate collaboration between Visual Tokens and Task Tokens, and limited interpretability of task decisions. To address these challenges, we propose an Explainable Task-Oriented Token Communication (ET-TokenCom) framework. By treating Tokens as unified units for information representation and transmission, the proposed framework constructs an end-to-end communication link that spans visual perception, wireless transmission, and task reasoning. At the transmitter, the ET-TokenCom framework extracts Visual Tokens from images to preserve low-level visual information. Meanwhile, Task Tokens generated by the FM are introduced to represent the target information and decision intent required by the current task. A Cross-Modal Attention (CMA) fusion mechanism is further designed, enabling Task Tokens to explicitly guide the selection, weighting, and transmission of Visual Tokens. At the receiver, the framework integrates Token decoding with an explainable output mechanism, where attention heatmaps are generated to highlight critical perceptual regions under different task objectives and reveal the influence of Task Tokens on the outputs. Finally, simulation results validate the effectiveness and robustness of the proposed ET-TokenCom framework.

任务导向通信可解释AI视觉令牌6G网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。