提出三级学习架构,让无人机群在搜救中自主协作并自我调节。
Intelligent Three Level Learning Architecture for Autonomous UAV Swarms in Search and Rescue
- 分层融合神经可塑性、多智能体强化学习与元学习,对应反射、技能、决策层级。
- 满足安全、最优、不饿死等六类形式化保障,动态下新增韧性与渐进改进能力。
- 适合复杂救援场景,为自主无人机群设计提供理论框架与实践路径。
本文提出一种新型三级分层学习架构,用于执行搜救任务的自主无人机群。与传统在各层级采用单一学习范式的方法不同,该架构集成三种质异的学习机制:个体适应采用赫布神经可塑性,战术协调结合图神经网络与行为树的多智能体强化学习,战略决策则融合模型无关元学习、BDI推理与数字孪生。架构通过六个组件(BDI、行为树、GNN、MARL、神经可塑性、元学习)中的22个架构契约进行形式化,共同提供六类形式化保证:安全性、预算正确性、最优性、活性、无饥饿性及跨层级一致性。引入“集群元认知”作为三层次结构化交互产生的组合性质,使群组可监控自身认知状态并切换策略。五种针对搜救任务类型的构造性进展函数连接抽象优化理论与具体操作场景。主集成定理表明,当所有契约满足时,混合神经符号系统保持全部六类保障。针对动态学习情况,新增五个契约,扩展出认知韧性、优雅退化与单调元改进三项新保障。理论分析证明该架构克服了现有分层强化学习的五大根本局限。
原文摘要 · Abstract (English)
This paper presents a novel three level hierarchical learning architecture for autonomous UAV swarms performing search and rescue operations. Unlike conventional approaches that apply a single learning paradigm across all hierarchy levels, the proposed architecture integrates three qualitatively different learning mechanisms corresponding to the biological hierarchy of reflexes, skills, and reasoning such as Hebbian neuroplasticity for individual agent adaptation, multi agent reinforcement learning with graph neural networks and behavior trees for tactical coordination, and model agnostic meta learning with BDI reasoning and a digital twin for strategic decision making. The architecture is formalized through twenty two architectural contracts organized across six components such as BDI, Behavior Trees, GNN, MARL, Neuroplasticity, Meta Learning that collectively provide six classes of formal guarantees such as safety, budget correctness, optimality, liveness, starvation freedom, and inter level consistency. We introduce Swarm Meta Cognition as a compositional property arising from the structured interaction of all three levels, enabling the swarm to monitor its own cognitive state and switch between cognitive strategies. Five constructive progress functions for SAR task types bridge the gap between abstract optimization theory and concrete operational scenarios. The main integration theorem establishes that when all contracts are satisfied, the hybrid neuro-symbolic system preserves all six guarantee classes. For the dynamic case with active learning, five new contracts extend the framework with three additional guarantees such as cognitive resilience, graceful degradation, and monotonic meta improvement. Theoretical analysis demonstrates that the architecture addresses five fundamental limitations of existing hierarchical RL approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。