arXiv:2509.23590eess.IVeess.SP2025-09

基于基础模型的自适应语义图像传输,提升动态无线环境下的传输鲁棒性。

Foundation Model-Based Adaptive Semantic Image Transmission for Dynamic Wireless Environments

  • 按任务需求分解图像语义,优先传输关键对象与细节。
  • 在BDD100K数据集上,SSIM、FID等指标优于现有方法。
  • 适合自动驾驶等动态无线场景,兼具高效与抗干扰能力。

基于基础模型的语义传输在无线图像通信中展现出巨大潜力。然而,现有方法存在两大局限:(i) 忽视特定下游任务中语义成分的重要性差异;(ii) 未能充分融合无线领域知识,导致在动态信道下鲁棒性不足。为此,本文提出一种面向动态无线环境(如自动驾驶)的基础模型自适应语义图像传输系统。该系统将图像分解为语义分割图与压缩表示,实现对关键物体和细粒度纹理的任务感知优先传输。通过任务自适应预编码机制,根据提取特征的语义重要性分配无线资源。为保障预编码所需的精确信道信息,采用条件扩散模型构建信道估计知识图(CEKM),融合用户位置、速度及稀疏信道样本,训练场景专用轻量级估计算法。接收端利用条件扩散模型从接收的语义特征重建高质量图像,有效应对信道损伤与部分数据丢失。基于QuaDRiGa生成多场景信道,在BDD100K数据集上的仿真结果表明,所提方法在感知质量(SSIM、LPIPS、FID)、任务精度(IoU)和传输效率方面均优于现有方法,验证了任务感知语义分解、场景自适应信道估计与扩散重建融合的有效性。

原文摘要 · Abstract (English)

Foundation model-based semantic transmission has recently shown great potential in wireless image communication. However, existing methods exhibit two major limitations: (i) they overlook the varying importance of semantic components for specific downstream tasks, and (ii) they insufficiently exploit wireless domain knowledge, resulting in limited robustness under dynamic channel conditions. To overcome these challenges, this paper proposes a foundation model-based adaptive semantic image transmission system for dynamic wireless environments, such as autonomous driving. The proposed system decomposes each image into a semantic segmentation map and a compressed representation, enabling task-aware prioritization of critical objects and fine-grained textures. A task-adaptive precoding mechanism then allocates radio resources according to the semantic importance of extracted features. To ensure accurate channel information for precoding, a channel estimation knowledge map (CEKM) is constructed using a conditional diffusion model that integrates user position, velocity, and sparse channel samples to train scenario-specific lightweight estimators. At the receiver, a conditional diffusion model reconstructs high-quality images from the received semantic features, ensuring robustness against channel impairments and partial data loss. Simulation results on the BDD100K dataset with multi-scenario channels generated by QuaDRiGa demonstrate that the proposed method outperforms existing approaches in terms of perceptual quality (SSIM, LPIPS, FID), task-specific accuracy (IoU), and transmission efficiency. These results highlight the effectiveness of integrating task-aware semantic decomposition, scenario-adaptive channel estimation, and diffusion-based reconstruction for robust semantic transmission in dynamic wireless environments.

语义传输扩散模型自适应通信自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。