arXiv:2607.15708cs.RO2026-07

无需绝对定位,多机器人通过视觉和通信自动生成相对位姿

Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations

论文配图:Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations
图 1 · 摘自论文原文
  • 用Transformer图网络隐式构建团队中心参考系,实现去中心化估计
  • 仿真训练后跨场景、跨队伍规模泛化,实测定位误差仅0.22米
  • 支持无单点故障,通信链路损失71%仍稳定运行,适合真实部署

传统领导者-跟随者编队控制存在单点故障和误差传播问题,且依赖不适用于无GPS环境的绝对定位传感器。本文提出一种基于视觉的端到端相对位姿估计算法,将每个机器人的单目图像与通信图中的消息直接映射为6自由度相对位姿。核心是隐式虚拟领航员(IVL):一个位于团队质心的非物理参考系,由基于Transformer的图神经网络隐式学习得到,使估计过程无特权节点且无需绝对定位。该方法同时输出校准后的随机不确定性(异方差GNLL)和认知不确定性(MC Dropout),在仿真与真实测试集上系统对比。仅在仿真中训练即可泛化至未见场景、更大队伍规模及外部真实基准。移除任一机器人最多导致中位数误差1.24倍,移除71%通信链路时位置误差仅增加1.77倍,无需重训练。在单一平台真实数据训练后,可直接迁移至异构团队,在实体机器人上实现0.22米、1.6°的相对位姿估计精度,并驱动闭环编队控制。

原文摘要 · Abstract (English)

Classical leader-follower formation control suffers from single points of failure and error propagation, and relies on absolute localization sensors that are ill-suited for GPS-denied environments. We present a learned, vision-only estimator that maps each robot's monocular image, together with messages exchanged over a communication graph, directly to its 6-DoF relative pose. Its key ingredient is the implicit virtual leader (IVL): a non-physical reference frame at the team centroid, implicitly learned inside a Transformer-based graph neural network, so that estimation has no privileged node and needs no absolute localization. The estimator additionally reports well-calibrated aleatoric (heteroscedastic GNLL) uncertainty alongside epistemic (MC~Dropout) uncertainty, compared systematically across simulation and real-world test sets. Trained only in simulation, the estimator generalizes to unseen scenes, to larger unseen team sizes, and to an external real-world benchmark. It exhibits no single point of failure: removing any one robot costs at most $1.24\times$ the median removal, and removing $71\%$ of the communication links costs $1.77\times$ in position error without retraining. Trained on real-robot data from a single platform, it transfers without modification to a heterogeneous team, estimating relative pose to $0.22$\,m and $1.6^\circ$ on physical robots, where it drives closed-loop formation control.

多机器人视觉定位去中心化图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。