让边缘视觉物联网按语义重要性传输图像,省带宽还保关键信息。
Semantic-Aware Generative Image Transmission for Resource-Constrained Visual IoT Systems

- 用语义重要性筛选要传的图像令牌,只发关键部分。
- 0.074bpp下用44.6%比特达29.9dB清晰度,比基准少传一半多。
- 适合低带宽、需保留目标物体的工业监控等场景。
资源受限的视觉物联网系统(如边缘摄像头、无人传感平台、工业检测节点和远程监测传感器)常需通过低速无线链路向边缘/云服务传输任务相关的视觉证据。现有图像通信方法通常压缩或传输完整的全局表示,难以利用接收端的生成恢复能力。本文提出一种面向边缘辅助视觉物联网的语义感知生成式图像传输框架。由物联网视觉传感器捕获的图像经VQ编码器编码为离散令牌网格。在物联网发射端或附近网关处,基于预测熵与局部结构复杂度估算的令牌可恢复性,与实例分割及类别感知评分获得的语义重要性融合,再由空间分散采样器在比特率预算内选择待传输令牌。发射端仅发送保留令牌的量化索引和二值掩码图,接收端通过采用哈顿序列调度的MaskGIT恢复被遮蔽的令牌。在Kodak和VisDrone场景下,于AWGN和瑞利信道中的实验表明,该方法为窄带视觉物联网链路提供了灵活的比特率-质量权衡。在0.074 bpp时,仅需0.167 bpp DeepJSCC/WITT参考方案44.6%的传输比特,即达29.9 dB PSNR。在Kodak上的伪真值下游检测实验进一步显示,在30%和50%掩码比率下,语义感知掩码比随机掩码更有效保留任务相关物体。
原文摘要 · Abstract (English)
Resource-constrained visual Internet of Things (IoT) systems, such as edge cameras, unmanned sensing platforms, industrial inspection nodes, and remote monitoring sensors, often need to transmit task-relevant visual evidence over low-rate wireless links to an edge/cloud service. Existing image communication methods usually compress or transmit complete global representations, leaving limited room to exploit receiver-side generative restoration. This paper proposes a semantic-aware generative image transmission framework for edge-assisted visual IoT. The image captured by an IoT visual sensor is encoded into a discrete token grid by a VQ encoder. At the IoT transmitter or nearby gateway, token recoverability, estimated from prediction entropy and local structure complexity, is fused with semantic importance obtained from instance segmentation and category-aware scoring. A spatial dispersal sampler then selects the tokens to be transmitted under a bitrate budget. The transmitter sends only the quantization indices of kept tokens and a binary mask map, while the edge/cloud receiver recovers masked tokens through MaskGIT with Halton sequence scheduling. Experiments on Kodak and VisDrone scenes under AWGN and Rayleigh channels show that the proposed method provides a flexible bitrate-quality tradeoff for narrowband visual IoT links. At 0.074 bpp, it uses 44.6% of the transmitted bits of the 0.167-bpp DeepJSCC/WITT reference while achieving 29.9 dB PSNR. A pseudo-GT downstream detection study on Kodak further shows that semantic-aware masking preserves task-relevant objects better than random masking at both 30% and 50% mask ratios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。