将符号强化学习融入无人机自主系统,提升动态任务规划能力。
Integrating Symbolic RL Planning into a BDI-based Autonomous UAV Framework: System Integration and SIL Validation
- 结合BDI架构与符号强化学习,实现动态任务规划。
- 仿真测试中任务效率提升75%,路径距离显著减少。
- 适合需要高可靠性决策的复杂无人机任务场景。
现代自主无人机任务日益需要能无缝整合结构化符号规划与自适应强化学习(RL)的软件框架。传统基于规则的架构虽具备稳健的结构化推理能力,但在动态复杂环境中难以满足自适应符号规划需求。符号强化学习(SRL)通过规划领域定义语言(PDDL)显式融合领域知识与操作约束,显著提升无人机决策的可靠性与安全性。本文提出AMAD-SRL框架,是针对无人机自主任务代理(AMAD)认知多智能体架构的扩展与优化,引入符号强化学习以实现动态任务规划与执行。在与预期硬件在环仿真(HILS)平台结构一致的软件在环(SIL)环境中验证了该框架。实验结果表明模块间集成稳定、互操作性强,可成功完成从BDI驱动到符号强化学习驱动规划阶段的过渡,且任务表现一致。具体评估了一个目标捕获场景:无人机先规划巡检路径,再动态调整进入路径以锁定目标并避开威胁区域。在该SIL评估中,相比基于覆盖率的基线,任务效率提升约75%,以行驶距离减少为衡量标准。本研究为处理复杂无人机任务奠定了坚实基础,并探讨了后续改进与验证方向。
原文摘要 · Abstract (English)
Modern autonomous drone missions increasingly require software frameworks capable of seamlessly integrating structured symbolic planning with adaptive reinforcement learning (RL). Although traditional rule-based architectures offer robust structured reasoning for drone autonomy, their capabilities fall short in dynamically complex operational environments that require adaptive symbolic planning. Symbolic RL (SRL), using the Planning Domain Definition Language (PDDL), explicitly integrates domain-specific knowledge and operational constraints, significantly improving the reliability and safety of unmanned aerial vehicle (UAV) decision making. In this study, we propose the AMAD-SRL framework, an extended and refined version of the Autonomous Mission Agents for Drones (AMAD) cognitive multi-agent architecture, enhanced with symbolic reinforcement learning for dynamic mission planning and execution. We validated our framework in a Software-in-the-Loop (SIL) environment structured identically to an intended Hardware-In-the-Loop Simulation (HILS) platform, ensuring seamless transition to real hardware. Experimental results demonstrate stable integration and interoperability of modules, successful transitions between BDI-driven and symbolic RL-driven planning phases, and consistent mission performance. Specifically, we evaluate a target acquisition scenario in which the UAV plans a surveillance path followed by a dynamic reentry path to secure the target while avoiding threat zones. In this SIL evaluation, mission efficiency improved by approximately 75% over a coverage-based baseline, measured by travel distance reduction. This study establishes a robust foundation for handling complex UAV missions and discusses directions for further enhancement and validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。