多智能体协同探索新场景,用混合方法提升效率与覆盖精度。
Hybrid Belief Reinforcement Learning for Efficient Coordinated Spatial Exploration
- 先用概率模型构建空间认知,再用强化学习优化路径规划。
- 相比基线,奖励提升10.8%,收敛速度加快38%。
- 适合需高效协作探测的无人机服务任务。
多个自主智能体在空间异质需求下协同探索与服务,需联合学习未知空间模式并规划最优轨迹以提升任务性能。纯模型驱动方法虽能提供结构化不确定性估计,但缺乏自适应策略学习;而深度强化学习在无空间先验时样本效率差。本文提出一种混合信念强化学习(HBRL)框架:第一阶段,智能体利用对数高斯泊松过程(LGCP)构建空间信念,并通过路径互信息(PathMI)规划器执行多步前瞻的信息驱动轨迹;第二阶段,将轨迹控制转移至软演员-评论家(SAC)智能体,通过双通道知识迁移实现热启动——信念状态初始化提供空间不确定性,回放缓冲区种子注入前期探索生成的示范轨迹。采用方差归一化的重叠惩罚机制,使共享信念状态实现协调覆盖,鼓励高不确定性区域合作感知,抑制已充分探索区域的重复覆盖。该框架在多无人机无线服务供给任务中评估,结果表明其累积奖励比基线高10.8%,收敛速度加快38%,消融实验验证双通道迁移优于单一通道。
原文摘要 · Abstract (English)
Coordinating multiple autonomous agents to explore and serve spatially heterogeneous demand requires jointly learning unknown spatial patterns and planning trajectories that maximize task performance. Pure model-based approaches provide structured uncertainty estimates but lack adaptive policy learning, while deep reinforcement learning often suffers from poor sample efficiency when spatial priors are absent. This paper presents a hybrid belief-reinforcement learning (HBRL) framework to address this gap. In the first phase, agents construct spatial beliefs using a Log-Gaussian Cox Process (LGCP) and execute information-driven trajectories guided by a Pathwise Mutual Information (PathMI) planner with multi-step lookahead. In the second phase, trajectory control is transferred to a Soft Actor-Critic (SAC) agent, warm-started through dual-channel knowledge transfer: belief state initialization supplies spatial uncertainty, and replay buffer seeding provides demonstration trajectories generated during LGCP exploration. A variance-normalized overlap penalty enables coordinated coverage through shared belief state, permitting cooperative sensing in high-uncertainty regions while discouraging redundant coverage in well-explored areas. The framework is evaluated on a multi-UAV wireless service provisioning task. Results show 10.8% higher cumulative reward and 38% faster convergence over baselines, with ablation studies confirming that dual-channel transfer outperforms either channel alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。