arXiv:2504.15876cs.ROcs.AI2025-04被引 1

双向强化学习让机器人团队高效应对对抗场景

Bidirectional Task-Motion Planning Based on Hierarchical Reinforcement Learning for Strategic Confrontation

  • 分层强化学习实现任务与运动的双向交互
  • 对抗胜率超80%,决策时间低于0.01秒
  • 适合大规模机器人协同与真实场景部署

在群体机器人对抗场景中,战略对抗需要高效融合离散指令与连续动作的决策。传统任务与运动规划方法采用单向分层结构,难以捕捉两层间的依赖关系,限制了动态环境中的适应性。本文提出一种基于分层强化学习的新型双向方法,实现两层间的动态交互。该方法有效将指令映射为任务分配,将动作映射为路径规划,并通过交叉训练技术增强分层框架内的学习效果。此外,引入轨迹预测模型,弥合抽象任务表示与可执行规划目标之间的鸿沟。实验表明,该方法在对抗中胜率超过80%,决策时间低于0.01秒,优于现有方法。大规模测试与真实机器人实验进一步验证了其泛化能力与实际应用价值。

原文摘要 · Abstract (English)

In swarm robotics, confrontation scenarios, including strategic confrontations, require efficient decision-making that integrates discrete commands and continuous actions. Traditional task and motion planning methods separate decision-making into two layers, but their unidirectional structure fails to capture the interdependence between these layers, limiting adaptability in dynamic environments. Here, we propose a novel bidirectional approach based on hierarchical reinforcement learning, enabling dynamic interaction between the layers. This method effectively maps commands to task allocation and actions to path planning, while leveraging cross-training techniques to enhance learning across the hierarchical framework. Furthermore, we introduce a trajectory prediction model that bridges abstract task representations with actionable planning goals. In our experiments, it achieves over 80% in confrontation win rate and under 0.01 seconds in decision time, outperforming existing approaches. Demonstrations through large-scale tests and real-world robot experiments further emphasize the generalization capabilities and practical applicability of our method.

群体机器人强化学习任务规划对抗策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。