用视觉无线联合自监督学习,让模型跨场景稳定识别环境
CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications
- 用软对比对齐损失融合图像与无线信号特征
- 跨场景迁移准确率达77.38%,提升4.61个百分点
- 适合6G无线感知与多模态感知研究者
同步摄像头与无线测量捕捉同一场景,但视角、流量、光照和传播几何变化时,原有表征常失效。本文提出CM-MAE,一种面向视觉-无线应用的自监督预训练框架。模型仅使用RGB帧和DeepSense 6G中的64波束接收功率向量进行预训练,不依赖射线追踪路径、校准深度或波束索引标签。核心采用软对比对齐损失,基于波束功率分布相似性构建目标分布,避免方向响应相似但非相同样本被误判为负例。掩码联合解码器通过模态丢弃重建隐藏视觉块与无线角度簇,提供局部重建目标。微调阶段采用差分率策略,新融合头快速适应而编码器缓慢更新。在序列无关的DeepSense 6G协议下,引入软对齐损失使线性探测平均准确率从24.88%提升至29.49%。轻度融合微调在未见场景6–8上达77.38% Top-1准确率,可选归纳归一化适配可达78.69%。因推理时使用实时64波束功率向量,结果应视为表征迁移诊断,而非主动波束预测或降采样宣称。
原文摘要 · Abstract (English)
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。