提出分层人机协同强化学习框架,提升无人机对抗任务的训练效率与性能。
A Systematic Approach to Design Real-World Human-in-the-Loop Deep Reinforcement Learning: Salient Features, Challenges and Trade-offs
- 构建三层学习架构:自学习、模仿学习与迁移学习协同工作。
- 实测显示人机协作使训练更快、性能更高,且建议适度提供指导信息。
- 适合研究人机协作、复杂决策系统或实际部署强化学习的开发者参考。
随着深度强化学习(DRL)的普及,人机协同(HITL)方法有望革新决策问题解决方式,推动人与AI的合作。本文提出一种多层分层式HITL DRL算法,包含自学习、模仿学习和迁移学习三类学习机制,并支持奖励、动作与示范三种人类输入形式。文章系统分析了在复杂问题中实施HITL所面临的挑战、权衡与优势。通过真实世界无人飞行器(UAV)场景验证:敌方无人机群攻击受限区域,目标是设计可扩展的HITL DRL算法,使己方无人机在敌方抵达前完成拦截。采用获奖开源工具Cogment实现方案,实验表明:(a) HITL显著加速训练并提升性能;(b) 人类建议作为梯度方法的引导方向,降低方差;(c) 建议量需适中,过多或过少均会导致过拟合或欠拟合。最后,展示了人机协作在应对过载攻击与诱饵攻击两种真实复杂场景中的作用。
原文摘要 · Abstract (English)
With the growing popularity of deep reinforcement learning (DRL), human-in-the-loop (HITL) approach has the potential to revolutionize the way we approach decision-making problems and create new opportunities for human-AI collaboration. In this article, we introduce a novel multi-layered hierarchical HITL DRL algorithm that comprises three types of learning: self learning, imitation learning and transfer learning. In addition, we consider three forms of human inputs: reward, action and demonstration. Furthermore, we discuss main challenges, trade-offs and advantages of HITL in solving complex problems and how human information can be integrated in the AI solution systematically. To verify our technical results, we present a real-world unmanned aerial vehicles (UAV) problem wherein a number of enemy drones attack a restricted area. The objective is to design a scalable HITL DRL algorithm for ally drones to neutralize the enemy drones before they reach the area. To this end, we first implement our solution using an award-winning open-source HITL software called Cogment. We then demonstrate several interesting results such as (a) HITL leads to faster training and higher performance, (b) advice acts as a guiding direction for gradient methods and lowers variance, and (c) the amount of advice should neither be too large nor too small to avoid over-training and under-training. Finally, we illustrate the role of human-AI cooperation in solving two real-world complex scenarios, i.e., overloaded and decoy attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。