用真实数据研究机器学习预测LoRa路径损耗,发现样本量越大效果越好。
A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN

- 用随机森林和最近邻算法结合地形与坐标数据预测路径损耗。
- 最大训练集下模型RMSE低于6.5 dB,优于基准模型的9.7 dB。
- 模型对新网关位置泛化能力有限,尤其坐标仅靠位置的k-NN表现差。
低功耗广域网络(如LoRa)在智慧城市建设中日益普及,准确的路径损耗预测对网络规划至关重要。传统经验传播模型在此类场景中精度有限。本文系统分析了机器学习模型在LoRa路径损耗预测中的表现,基于城市部署的真实测量数据,考察了训练集规模对预测精度的影响。方法采用融合激光雷达地形特征的随机森林(RF)和仅使用坐标的k-近邻(k-NN),并与经典经验模型及专用LPWAN模型对比。在随机池化分割下,两种机器学习模型在所有训练集规模下均优于基线模型;在最大训练规模时,其均方根误差(RMSE)低于6.5 dB,而最佳基线为9.7 dB,表明模型在部署范围内具有高精度插值能力。通过留一网关验证进一步检验:随机森林在部分未见网关上出现轻微退化,但个别网关误差显著上升;而仅依赖坐标的k-NN在未见位置上性能大幅下降。
原文摘要 · Abstract (English)
Low Power Wide Area Networks like LoRa are increasingly deployed for smart city applications, requiring accurate path loss prediction for effective network planning. Traditional (empirical) propagation models often exhibit limited accuracy in these scenarios. We investigate machine learning models for LoRa path loss prediction, systematically analyzing how prediction accuracy scales with training set size using real-world measurements from an urban deployment. Our approach employs a Random Forest with LiDAR-derived terrain features and k-Nearest Neighbors with coordinate data, comparing their performance against established empirical models and specialized LPWAN models. Under random pooled splits, both ML models consistently outperform the considered baseline models across the evaluated training-set sizes. At maximum training size, they achieve RMSE values below 6.5 dB compared to 9.7 dB for the best baseline, indicating accurate within-deployment interpolation. A leave-one-gateway-out check qualifies this result: RF shows placement-dependent transfer to held-out gateways, with moderate degradation for several gateways but larger errors for others, whereas coordinate-only k-NN degrades substantially when the gateway location is unseen
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。