arXiv:2508.03736cs.CVcs.AI2025-08被引 1

用视觉变压器融合无线信号与地图数据,提升智慧城市建筑映射精度。

Fusion of Pervasive RF Data with Spatial Images via Vision Transformers for Enhanced Mapping in Smart Cities

  • 用视觉变压器统一处理无线信号和地图数据,捕捉空间依赖关系。
  • 合成数据上达65.3%宏观交并比,显著优于仅用无线信号或地图的基线。
  • 方法适合智能城市高精度地图构建,可扩展至大区域部署。

本文提出一种基于深度学习的方法,利用DINOv2架构,将开源平台可能出错的地图与多设备采集的无线射频(RF)数据融合,以提升建筑映射精度。不同于以往方法,该方法采用视觉变压器架构,在统一框架中联合处理RF与地图模态,有效捕捉空间依赖性和结构先验。评估使用华为共同生成的合成数据集,并对RF数据添加可控噪声以模拟真实场景。此外,训练了仅依赖聚合路径损耗信息的模型。通过交并比(IoU)、豪斯多夫距离和钱弗距离三个指标评估性能。结果表明,该方法在合成数据上实现65.3%的宏观IoU,显著高于错误地图基线(40.1%)、文献中的纯RF方法(37.3%)以及自设计的非AI融合基线(42.2%)。在奥斯陆真实数据上验证,最优融合模型达64.9%宏观IoU。还提出通过重叠窗口分块策略实现大区域部署。

原文摘要 · Abstract (English)

In this paper, we present a deep learning-based approach that integrates the DINOv2 architecture to improve building mapping by combining (possibly erroneous) maps from open-source platforms with pervasive radio frequency (RF) data collected from multiple wireless user equipments and base stations. Unlike prior methods, our approach leverages a vision transformer-based architecture to jointly process both RF and map modalities within a unified framework, effectively capturing spatial dependencies and structural priors for enhanced mapping accuracy. For the evaluation purposes, we employ a synthetic dataset co-produced by Huawei. To address the challenges associated with real-world data imperfections, we introduce controlled noise to its RF data so as to simulate real-world conditions. Additionally, we develop and train a model that leverages only aggregated path loss information to tackle the mapping problem. We measure the results according to three performance metrics: the Jaccard index (intersection over union, IoU), the Hausdorff distance, and the Chamfer distance. Our design achieves a macro IoU of 65.3%, significantly surpassing (i) the erroneous maps baseline, which yields 40.1%, (ii) an RF-only method from the literature, which yields 37.3%, and (iii) a non-AI fusion baseline that we designed which yields 42.2%. The comparative evaluation highlights the limitations of relying solely on RF data or on spatial data, as well as the effectiveness that AI can have on fusing data towards enhancing smart city mapping accuracy. We further validate our method on real-world data from the Oslo region, complementing the synthetic evaluation with a real deployment setting, where our best fusion model reaches 64.9% macro IoU. We additionally outline a strategy for deploying the model over larger areas by tiling the region with overlapping windows.

智慧城市场景多模态融合视觉变压器射频数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。