arXiv:2609.05515cs.RO2026-09

多机器人用时空高斯过程实现高效目标监控,降低不确定性20%。

Multi-robot Learning-based Informative Path Planning Using Spatio-Temporal Gaussian Process Kalman Filter

论文配图:Multi-robot Learning-based Informative Path Planning Using Spatio-Temporal Gaussian Process Kalman Filter
图 1 · 摘自论文原文
  • 将目标位置建模为网格上的潜在场,支持任意视场和距离噪声。
  • 递归更新考虑目标移动与信息过时,使不确定性随时间增长。
  • 分布式部署,仅交换压缩信念信息,适合真实无人机搜索场景。

多机器人持续目标监测中的信息路径规划需同时处理空间不确定性、时间演化及传感通信约束。现有基于学习的多机器人方法虽使用高斯过程(GPs)表示目标不确定性,但常依赖简化传感模型且采用集中式信念更新。本文提出一种基于网格的时空高斯过程-卡尔曼滤波框架,不再为每个目标维护独立的高斯过程,而是将匿名目标存在性建模为离散工作区网格上的单一潜在场。所提递归更新考虑相机视野内所有可见单元,支持任意视场与距离相关的噪声。基于高斯过程一致性的时序更新通过随时间膨胀不确定性来应对移动目标与过时信息。为实现去中心化部署,每台机器人维护自身映射器,并交换紧凑的信念摘要而非原始测量值;接收的信念通过对角协方差交集融合,在未知机器人间相关性下保持保守性。将映射器与基于图的邻近选择强化学习策略集成。仿真结果显示,平均目标不确定性比基于学习及经典拍卖/覆盖基线低约20%,目标访问率更优;两架无人机的真实世界实验成功实现超过7000平方米户外区域的多机器人搜索任务迁移。

原文摘要 · Abstract (English)

Multi-robot informative path planning (IPP) for persistent target monitoring requires robots to reason about spatial uncertainty, temporal evolution, and practical sensing and communication constraints. Recent learning-based multi-robot IPP methods use Gaussian Processes (GPs) for target uncertainty, but often rely on simplified sensing models and centralized belief updates. We propose a grid-based spatio-temporal GP-Kalman filtering framework for learning-based multi-robot IPP. Instead of maintaining one GP per target, we represent anonymous target presence as a single latent field over a discrete workspace grid. The proposed recursive update considers all visible cells inside a camera footprint and supports arbitrary fields of view and range-dependent noise. A GP-consistent temporal process update accounts for moving targets and stale information by inflating uncertainty over time. For decentralized deployment, each robot maintains its own mapper and exchanges compact belief summaries rather than raw measurements. Received beliefs are fused using diagonal covariance intersection to remain conservative under unknown inter-robot correlations. We integrate the mapper with a reinforcement-learning policy for graph-based neighbor selection. Simulation benchmarks show about 20% lower average target uncertainty and improved target visitation compared with learning-based and classical auction/coverage baselines. Real-world two-UAV experiments demonstrate transfer to outdoor multi-robot search over a large field of more than 7000 square meters.

多机器人路径规划高斯过程强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。