arXiv:2505.01870cs.ITeess.IV2025-05被引 5

ResiTok让图像在极低带宽下仍能清晰传输,还能自动适应网络波动。

ResiTok: A Resilient Tokenization-Enabled Framework for Ultra-Low-Rate and Robust Image Transmission

  • 将图像分层为关键令牌和补充细节令牌,支持渐进式编码
  • 在极低带宽下仍保持高重建质量,语义相似度显著领先
  • 适合弱网环境下的实时图像传输,如无人机、远程医疗

无线网络中实时传输视觉数据仍面临巨大挑战,尤其在带宽受限或连接薄弱的情况下。本文提出一种新型抗干扰分词框架ResiTok,专为超低速率图像传输设计,可在极端条件下实现卓越鲁棒性与高质量重建。通过将视觉信息重组为包含关键令牌和补充细节令牌的分层令牌组,实现渐进编码与平滑退化。核心创新在于结合零值训练策略的1D抗损分词方法,系统模拟训练中令牌丢失,使神经网络能从不完整令牌集重建图像。此外,自适应编码调制设计根据信道状况动态分配资源,确保在极低带宽比下仍具备优异语义保真度与结构一致性。实验表明,ResiTok在语义相似度与视觉质量上均优于现有方法,尤其在恶劣信道条件下优势明显。

原文摘要 · Abstract (English)

Real-time transmission of visual data over wireless networks remains highly challenging, even when leveraging advanced deep neural networks, particularly under severe channel conditions such as limited bandwidth and weak connectivity. In this paper, we propose a novel Resilient Tokenization-Enabled (ResiTok) framework designed for ultra-low-rate image transmission that achieves exceptional robustness while maintaining high reconstruction quality. By reorganizing visual information into hierarchical token groups consisting of essential key tokens and supplementary detail tokens, ResiTok enables progressive encoding and graceful degradation of visual quality under constrained channel conditions. A key contribution is our resilient 1D tokenization method integrated with a specialized zero-out training strategy, which systematically simulates token loss during training, empowering the neural network to effectively compress and reconstruct images from incomplete token sets. Furthermore, the channel-adaptive coding and modulation design dynamically allocates coding resources according to prevailing channel conditions, yielding superior semantic fidelity and structural consistency even at extremely low channel bandwidth ratios. Evaluation results demonstrate that ResiTok outperforms state-of-the-art methods in both semantic similarity and visual quality, with significant advantages under challenging channel conditions.

图像传输低速率鲁棒性分词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。