arXiv:2602.13606cs.NIcs.AI2026-02被引 5

用多模态融合预测毫米波波束,大幅降低车联网通信开销

Multi-Modal Sensing and Fusion in mmWave Beamforming for Connected Vehicles: A Transformer Based Framework

  • 通过多模态编码与交叉注意力融合感知信息
  • 预测前15个波束准确率达96.72%,功率损失仅0.77dB
  • 适合高动态车联网场景下的低延迟波束管理

毫米波(mmWave)通信利用波束成形技术应对固有的路径损耗问题,被视为满足车联网日益增长的高吞吐量和低延迟需求的关键技术。然而,在高度动态的车辆环境中,采用标准波束成形方法常导致高昂的波束训练开销及通信可用时长减少,主要源于导频信号交换和遍历式波束测量。为此,我们提出一种基于Transformer的多模态感知与融合学习框架,作为替代方案以降低此类开销。该框架首先通过模态专用编码器提取感知模态的代表性特征,再利用多头跨模态注意力学习不同模态间的依赖关系与相关性,进而融合多模态特征,预测出最优的top-k波束,从而主动建立最佳视距链路。为验证该框架的通用性,我们在四个真实世界多模态与60 GHz mmWave无线感知数据的车对基础设施(V2I)和车对车(V2V)场景中进行了全面实验。结果表明,所提框架(i)在预测前15个波束时准确率最高达96.72%;(ii)平均功率损失约为0.77 dB;(iii)相较于标准方法,整体延迟和波束搜索空间开销分别降低86.81%和76.56%。

原文摘要 · Abstract (English)

Millimeter wave (mmWave) communication, utilizing beamforming techniques to address the inherent path loss limitation, is considered as one of the key technologies to support ever increasing high throughput and low latency demands of connected vehicles. However, adopting standard defined beamforming approach in highly dynamic vehicular environments often incurs high beam training overheads and reduction in the available airtime for communications, which is mainly due to exchanging pilot signals and exhaustive beam measurements. To this end, we present a multi-modal sensing and fusion learning framework as a potential alternative solution to reduce such overheads. In this framework, we first extract the representative features from the sensing modalities by modality specific encoders, then, utilize multi-head cross-modal attention to learn dependencies and correlations between different modalities, and subsequently fuse the multimodal features to obtain predicted top-k beams so that the best line-of-sight links can be proactively established. To show the generalizability of the proposed framework, we perform a comprehensive experiment in four different vehicle-to-infrastructure (V2I) and vehicle-to-vehicle (V2V) scenarios from real world multimodal and 60 GHz mmWave wireless sensing data. The experiment reveals that the proposed framework (i) achieves up to 96.72% accuracy on predicting top-15 beams correctly, (ii) incurs roughly 0.77 dB average power loss, and (iii) improves the overall latency and beam searching space overheads by 86.81% and 76.56% respectively for top-15 beams compared to standard defined approach.

毫米波通信车联网多模态融合波束成形

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。