arXiv:2609.07853cs.LGcs.AI2026-09

用大模型先验实现低比特率下语义通信的高保真重建。

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

论文配图:Foundation Models for Generalizable Semantic and Goal-Oriented Communication
图 1 · 摘自论文原文
  • 利用视觉语言大模型选择关键语义锚点,只传输稀疏信息。
  • 在0.039比特/像素下保持语义相似度0.87-0.90,优于现有方法。
  • 适合6G低带宽场景,尤其对未见图像有强泛化能力。

语义与目标导向通信在6G中日益受到关注,但在严格速率限制下,泛化能力仍是主要短板。现有系统易过拟合训练数据,在极低比特率下性能急剧下降,因试图压缩完整信号。本文提出基础模型引导的语义与目标导向通信(FMSGOC),利用广泛的视觉-语言基础模型先验来缓解过拟合。通过将发送内容与重建方式解耦,视觉-语言基础模型在接收端选择并传输一组稀疏的语义锚点,而微调过的扩散模型则完成掩码区域的生成重建。实验表明,FMSGOC在CIFAR-10上达到0.039比特/像素,语义保真度(余弦相似度)达0.87–0.90,在ImageNet上对未见过输入仍保持0.83–0.86的相似度,感知相似度分别为0.1278(CIFAR-10)和0.1558(ImageNet),显著优于强基准模型。

原文摘要 · Abstract (English)

Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. We introduce Foundation Model-Guided Semantic and Goal-Oriented Communication (FMSGOC), a framework that uses broad visual-linguistic Foundation Model priors to mitigate overfitting. It further improves rate efficiency by concentrating bits on sparse, goal-aligned anchors and relying on generative foundation-model priors to reconstruct the masked regions. By decoupling what to send from how to reconstruct, a vision-language foundation model selects and transmits a sparse set of semantic anchors, while a pretrained diffusion model, fine-tuned for masked completion, reconstructs the image at the receiver. In our experiments, FMSGOC reaches 0.039 bits per pixel (BPP), maintains high semantic fidelity (cosine similarity 0.87-0.90 on CIFAR-10), remains robust on previously unseen inputs (0.83-0.86 on ImageNet), and shows good perceptual similarity (0.1278/0.1558, CIFAR-10/ImageNet), outperforming strong end-to-end baselines at lower bit rates.

语义通信6G基础模型低比特率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。