用石汤追踪框架构建强化学习训练环境,提升无人机自主搜寻跟踪能力。
Stone Soup Multi-Target Tracking Feature Extraction For Autonomous Search And Track In Deep Reinforcement Learning Environment
- 将石汤追踪器嵌入Gymnasium环境,作为强化学习的特征提取器。
- 在空域搜索与跟踪任务中,强化学习智能体性能优于传统策略。
- 支持多种神经网络架构,可快速配置用于复杂场景训练。
未来军事航空平台需管理异构传感器以获取战场信息,而机器学习尤其是深度强化学习(DRL)被视为有前景的解决方案,但依赖高保真训练环境与特征提取器。本文提出一种基于石汤(Stone Soup)追踪框架的深度强化学习训练方法,将其组件嵌入Gymnasium环境,实现稳定、可配置的追踪器部署,配合Stable Baselines3进行训练。在空域搜索与跟踪任务中,智能体通过石汤生成的航迹列表进行决策。实验采用三种神经网络架构,在模拟场景中验证了该方法的有效性,结果表明:在石汤与Gymnasium环境中训练的强化学习智能体性能超越简单搜索与跟踪策略。
原文摘要 · Abstract (English)
Management of sensing resources is a non-trivial problem for future military air assets with future systems deploying heterogeneous sensors to generate information of the battlespace. Machine learning techniques including deep reinforcement learning (DRL) have been identified as promising approaches, but require high-fidelity training environments and feature extractors to generate information for the agent. This paper presents a deep reinforcement learning training approach, utilising the Stone Soup tracking framework as a feature extractor to train an agent for a sensor management task. A general framework for embedding Stone Soup tracker components within a Gymnasium environment is presented, enabling fast and configurable tracker deployments for RL training using Stable Baselines3. The approach is demonstrated in a sensor management task where an agent is trained to search and track a region of airspace utilising track lists generated from Stone Soup trackers. A sample implementation using three neural network architectures in a search-and-track scenario demonstrates the approach and shows that RL agents can outperform simple sensor search and track policies when trained within the Gymnasium and Stone Soup environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。