arXiv:2508.17920eess.IVcs.MM2025-08被引 2

用提示词引导多光谱图像分割中的跨模态语义融合,提升性能。

Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation

  • 一模态特征作提示,引导另一模态学习互补语义。
  • 跨注意力与SE网络结合,有效融合多模态特征。
  • 性能超越基准方法,适合自动驾驶等实时场景。

多模态语义通信因能提升下游任务表现而受到广泛关注。其核心挑战在于如何有效融合不同模态的特征,这需要从各模态中提取丰富且多样化的语义表示。为此,我们提出 ProMSC-MIS:一种面向多光谱图像分割的提示词驱动多模态语义通信系统。具体地,设计了一种预训练算法,使某一模态的特征作为提示,引导另一模态的单模态语义编码器学习多样化、互补的语义表征。进一步引入一个融合模块,结合交叉注意力机制与挤压-激励(SE)网络,高效融合跨模态特征。仿真结果表明,ProMSC-MIS 在多种信道-源压缩水平下均显著优于基准方法,同时保持低计算复杂度与存储开销。该方案在自动驾驶、夜间监控等领域具有广泛应用潜力。

原文摘要 · Abstract (English)

Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the extraction of rich and diverse semantic representations from each modality. To this end, we propose ProMSC-MIS, a Prompt-based Multimodal Semantic Communication system for Multi-spectral Image Segmentation. Specifically, we propose a pre-training algorithm where features from one modality serve as prompts for another, guiding unimodal semantic encoders to learn diverse and complementary semantic representations. We further introduce a semantic fusion module that combines cross-attention mechanisms and squeeze-and-excitation (SE) networks to effectively fuse cross-modal features. Simulation results show that ProMSC-MIS significantly outperforms benchmark methods across various channel-source compression levels, while maintaining low computational complexity and storage overhead. Our scheme has great potential for applications such as autonomous driving and nighttime surveillance.

多模态图像分割提示词语义通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。