面向丢包信道的通用视觉协同感知新模型,提升信息传输效率与泛化能力。
SoM-MTM: Synesthesia of Machines (SoM)-Driven Masked Token Model for Cooperative Perception over Packet Loss Channel

- 基于机器联觉思想设计掩码令牌模型,增强感知上下文学习能力。
- 在多种任务上显著提升感知性能,尤其在未见场景中表现优异。
- 可即插即用,兼顾模型轻量与可扩展性,适合实际部署。
为支持下一代移动网络中大规模、异构的视觉协同感知(CP)需求,智能高效的传感数据传输成为关键挑战。在通信网络与代理型人工智能(AI)融合的背景下,现有研究倾向于采用端到端神经网络简化通信模块,已在协同感知中展现出潜力。然而,这些方法仍局限于特定信道模型、协作模式和感知任务,未能充分挖掘强大视觉处理方法的通用性优势。为此,本文提出一种由机器联觉(SoM)驱动的掩码令牌模型(SoM-MTM),作为通用视觉协同感知的即插即用范式。受MAE等掩码图像建模方法启发,该模型具备强大的感知上下文学习能力,可在丢包信道中恢复受损特征,从而提升信息承载效率。基于Swin Transformer架构,SoM-MTM进一步通过外部路由MoE机制嵌入先验掩码信息,在协作过程中最大限度修复并增强环境感知特征。综合实验表明,SoM-MTM能在多种任务上持续提升感知性能,尤其在未见场景中展现强泛化能力,同时保持较低的模型开销与良好可扩展性。
原文摘要 · Abstract (English)
To support the large-scale and heterogeneous visual cooperative perception (CP) demands in next-generation mobile networks, intelligent and efficient sensory data transmission is a critical challenge. Under the emerging convergence of communication networks and agentic artificial intelligence (AI), existing research emphasizes utilizing end-to-end neural networks to simplify communication modules, which has shown promising potential for CP. However, these studies are still limited to specific channel models, cooperation modes, and perception tasks, failing to fully leverage powerful visual processing approaches to enhance universality. To address this, we propose a Synesthesia of Machines (SoM)-driven Masked Token Model, referred to as SoM-MTM, as a plug-and-play paradigm for generic visual CP. Inspired by masked image modeling methods such as MAE, it possesses great perceptual context learning capabilities to recover distorted features over packet loss channels, thereby improving information carrying efficiency. Building upon Swin Transformer, SoM-MTM further embeds prior masked information through an External Routing MoE mechanism, maximally repairing and enhancing environmental perception features during cooperation. Comprehensive experimental results confirm that SoM-MTM can consistently enhance perception performances on various tasks, especially strong generalization to unseen scenarios, while maintaining favorable model cost and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。