arXiv:2511.11190cs.LG2025-11

用强化学习提升蓝牙标签定位效率,抗干扰能力强。

LoRaCompass: Robust Reinforcement Learning to Efficiently Search for a LoRa Tag

  • 基于信号强度构建空间特征,通过策略蒸馏增强鲁棒性。
  • 在未知环境中定位成功率超90%,比现有方法快40%。
  • 适合移动传感器在复杂环境下快速精准搜寻标签。

长距离(LoRa)协议因覆盖范围广、功耗低,被广泛用于精神障碍者等易走失人群佩戴的标签。本文研究移动传感器在未知环境中,通过接收信号强度指示(RSSI)以最少移动次数(跳数)定位周期性广播的LoRa标签的序列决策问题。现有强化学习方法易受环境变化和信号波动影响,导致决策误差累积,定位精度下降。为此,我们提出LoRaCompass,一种具备强鲁棒性和高效率的强化学习模型。该模型通过空间感知特征提取器与策略蒸馏损失函数,从RSSI中学习鲁棒的空间表示,最大化靠近标签的概率;并引入受上限置信区间(UCB)启发的探索函数,逐步提升定位信心。我们在超过80km²的多种未见环境中验证了该方法,包括地面与无人机辅助场景。结果表明,其定位成功率超过90%(在100米范围内),相比现有方法提升40%;搜索路径长度(以跳数计)与初始距离呈线性关系,表现出优异的效率。

原文摘要 · Abstract (English)

The Long-Range (LoRa) protocol, known for its extensive range and low power, has increasingly been adopted in tags worn by mentally incapacitated persons (MIPs) and others at risk of going missing. We study the sequential decision-making process for a mobile sensor to locate a periodically broadcasting LoRa tag with the fewest moves (hops) in general, unknown environments, guided by the received signal strength indicator (RSSI). While existing methods leverage reinforcement learning for search, they remain vulnerable to domain shift and signal fluctuation, resulting in cascading decision errors that culminate in substantial localization inaccuracies. To bridge this gap, we propose LoRaCompass, a reinforcement learning model designed to achieve robust and efficient search for a LoRa tag. For exploitation under domain shift and signal fluctuation, LoRaCompass learns a robust spatial representation from RSSI to maximize the probability of moving closer to a tag, via a spatially-aware feature extractor and a policy distillation loss function. It further introduces an exploration function inspired by the upper confidence bound (UCB) that guides the sensor toward the tag with increasing confidence. We have validated LoRaCompass in ground-based and drone-assisted scenarios within diverse unseen environments covering an area of over 80km^2. It has demonstrated high success rate (>90%) in locating the tag within 100m proximity (a 40% improvement over existing methods) and high efficiency with a search path length (in hops) that scales linearly with the initial distance.

强化学习定位系统物联网鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。