用规则指导强化学习,让无人机在少训练下更安全高效完成搜救任务。
Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training
- 高层用预设规则提供可解释的行动建议与避险指引
- 低层强化学习在线适应,碰撞终止减少60%以上
- 适合对安全性要求高、数据有限的无人机实际部署
本文针对搜救无人机任务中仿真训练资源有限的问题,提出一种分层决策框架。该框架将离线生成的规则化高层指导器与在线学习的目标条件强化学习控制器结合。高层指导器基于结构化任务描述生成确定性规则,提供可解释的行动推荐、避让动作及场景依赖的仲裁权重。底层控制器通过任务定义的密集奖励在线学习,并利用融合规则元信息的模式感知优先级重放机制复用经验。在电池敏感多目标配送和复杂障碍环境中的移动目标配送两个任务上评估,所提方法显著提升早期安全性和样本效率,主要通过减少碰撞终止实现,同时保持对场景动态的在线适应能力。
原文摘要 · Abstract (English)
This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios under limited simulation training. The framework combines a fixed rule-based high-level advisor with an online goal-conditioned low-level reinforcement learning (RL) controller. To stress-test early adaptation, we also consider a strict no-pretraining deployment regime. The high-level advisor is defined offline from a structured task specification and compiled into deterministic rules. It provides interpretable mission- and safety-aware guidance through recommended actions, avoided actions, and regime-dependent arbitration weights. The low-level controller learns online from task-defined dense rewards and reuses experience through a mode-aware prioritized replay mechanism augmented with rule-derived metadata. We evaluate the framework on two tasks: battery-aware multi-goal delivery and moving-target delivery in obstacle-rich environments. Across both tasks, the proposed method improves early safety and sample efficiency primarily by reducing collision terminations, while preserving the ability to adapt online to scenario-specific dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。