用希尔伯特空间填充提升多机器人探索效率,减少重复覆盖。
Hilbert-Augmented Reinforcement Learning for Scalable Multi-Robot Coverage and Exploration
- 将希尔伯特空间索引融入DQN/PPO,指导机器人分布探索。
- 在网格覆盖任务中,覆盖率提升23%,冗余降低31%。
- 适用于资源受限的轮式/足式机器人,实现在室内可靠运行。
本文提出一种集成希尔伯特空间填充先验的覆盖框架,用于去中心化多机器人学习与执行。通过在DQN和PPO中引入基于希尔伯特的空间索引,结构化探索过程,减少稀疏奖励环境下的冗余行为,并在多机器人网格覆盖任务中验证了其可扩展性。进一步设计了一种航点接口,将希尔伯特序转换为曲率受限、时间参数化的SE(2)轨迹(平面运动:x, y, θ),实现资源受限机器人上的在线可行性。实验表明,相较于基线方法,该方法在覆盖率、冗余控制和收敛速度上均有显著提升。此外,在波士顿动力Spot足式机器人上进行了验证,于室内环境中成功执行生成轨迹,实现了高效且低冗余的覆盖。结果表明,几何先验能有效提升群体与足式机器人的自主性与可扩展性。
原文摘要 · Abstract (English)
We present a coverage framework that integrates Hilbert space-filling priors into decentralized multi-robot learning and execution. We augment DQN and PPO with Hilbert-based spatial indices to structure exploration and reduce redundancy in sparse-reward environments, and we evaluate scalability in multi-robot grid coverage. We further describe a waypoint interface that converts Hilbert orderings into curvature-bounded, time-parameterized SE(2) trajectories (planar (x, y, θ)), enabling onboard feasibility on resource-constrained robots. Experiments show improvements in coverage efficiency, redundancy, and convergence speed over DQN/PPO baselines. In addition, we validate the approach on a Boston Dynamics Spot legged robot, executing the generated trajectories in indoor environments and observing reliable coverage with low redundancy. These results indicate that geometric priors improve autonomy and scalability for swarm and legged robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。