用视觉辅助解决无线感知中多目标混淆和分辨率低的问题
Resolving Multi-Target Association in OFDM-based ISAC via Vision-aided Multi-Modal Learning

- 融合无线信号与车载图像,通过深度联合编码传输视觉信息
- 实现16厘米定位误差和10.8纳秒延迟误差,显著优于单模态系统
- 适合自动驾驶、智能交通等需要高精度多目标感知的场景
基于正交频分复用(OFDM)的集成感知与通信(ISAC)系统通常通过搜索延迟-多普勒图(DDM)中的峰值来提取目标参数。在多目标场景下,这会产生歧义:DDM无法揭示哪个物理目标对应哪个峰值,且处于同一延迟-多普勒分辨单元内的两个目标无法分离。本文提出一种视觉辅助的OFDM-ISAC框架,通过融合无线与视觉模态来解决上述问题。发射端使用深度联合源信道编码(DeepJSCC)对车载街景图像进行编码,并通过相同的OFDM波形传输;接收端重建图像后运行微调后的YOLOv5检测器,将每个目标的特征(边界框坐标和类别标签)与DDM及收发端几何关系通过学习型多模态网络融合。为稳定高维延迟与多普勒分类器的训练,引入以真实值所在区间为中心的三角形软标签的KL散度损失。在Blender渲染的车辆测试平台上,该框架实现了16厘米定位均方根误差(RMSE)和10.8纳秒延迟RMSE。消融实验表明,移除视觉模态会导致定位性能下降60倍。结果凸显了视觉在突破单模态ISAC的数据关联与分辨率瓶颈方面的潜力。
原文摘要 · Abstract (English)
Orthogonal frequency division multiplexing (OFDM)-based integrated sensing and communication (ISAC) systems commonly extract target parameters by peak-searching a delay-Doppler map (DDM) constructed from reflected pilots. In multi-target scenarios, this results in ambiguity: the DDM does not reveal which physical target produced which peak, and two targets within the same delay-Doppler resolution cell cannot be separated. We propose a vision-assisted OFDM-ISAC framework that resolves both limitations by fusing wireless and visual modalities. The transmitter encodes an onboard street-view image with deep joint source-channel coding (DeepJSCC) and transmits it over the same OFDM waveform used for sensing; the receiver reconstructs the image, runs a fine-tuned YOLOv5 detector and fuses the resulting per-target features (bounding-box coordinates and class labels) with the DDM and transmitter-receiver geometry through a learned multi-modal network. To stabilize training of the high dimensional delay and Doppler classifiers, we introduce a Kullback Leibler loss against triangular soft labels centered on the ground-truth bin. On a Blender-rendered vehicular testbed, the proposed framework achieves a 16 cm localization root mean square error (RMSE) and a 10.8 ns delay RMSE. An ablation study confirms that removing the visual modality causes a 60x degradation in localization. These results highlight the potential of vision to overcome the data-association and resolution limits of single-modality ISAC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。