用强化学习选海事监控传感器,省算力还接近全开效果。
Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

- 用强化学习每步只选一个传感器,不全开也不算太贵。
- 性能接近全传感器开启,但只激活一个,省资源。
- 适合实时海事追踪,尤其传感器种类多时。
本文提出一种基于信息增益的强化学习传感器选择框架,用于异构海事传感网络中的单船追踪。该方法受信息论启发:避免激活所有传感器或进行高成本的在线期望信息增益计算,转而通过学习到的策略在每个决策时刻选择一个与追踪相关性高的传感器。采用贝叶斯序贯蒙特卡洛追踪器从噪声测量中估计船舶状态,并提供非线性、非高斯条件下的信念表示。使用近端策略优化(Proximal Policy Optimization)代理,在塞浦路斯阿伊亚纳帕港的CMMI智能码头测试场地理参考仿真环境中,从五个部署传感器中选择最优。代理观测信念状态、检测历史、覆盖范围、传感器几何和实际信息增益特征。奖励定义为经可观测性掩码加权的实际信息增益项。最终测试模拟对比了该框架与随机单传感器选择、所有传感器同时开启的全开模式,以及我们之前提出的期望信息增益基线。结果表明,所提学习策略在仅每步激活一个传感器的情况下,实现了接近全开感知的追踪性能,且避免了期望信息增益选择所需的高成本在线熵搜索。
原文摘要 · Abstract (English)
This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch. A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors deployed in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with random single-sensor selection, always-on sensing using all sensors simultaneously, and the expected-information-gain sensor-selection baseline proposed in our previous work. Results show that the learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step and avoiding the computationally expensive online entropy search required by expected-information-gain selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。