arXiv:2603.12908cs.RO2026-03

多无人机协同定位目标,无需训练即可识别新物体。

GoalSwarm: Multi-UAV Semantic Coordination for Open-Vocabulary Object Navigation

  • 每架无人机用轻量2D地图共享语义信息,避免复杂3D建模。
  • 利用视觉大模型实现零样本识别,可定位未见过的物体。
  • 通过竞争性投标与路径优化,减少重复探索,适合复杂环境导航。

协同视觉语义导航是飞行机器人团队在未知环境中运行的基础能力。然而,由于机载部署重型感知模型的计算限制以及去中心化多智能体协调的复杂性,实现稳健的开放词汇物体目标导航仍具挑战。本文提出GoalSwarm,一种完全去中心化的多无人机框架,支持零样本语义物体目标导航。每架无人机通过从空中视角投影深度观测,协作构建轻量级2D俯视语义占用地图,消除全3D表示的计算负担,同时保留关键几何与语义结构。核心贡献有三:(1)集成零样本基础模型SAM3,实现开放词汇检测与像素级分割,无需任务特定训练即可识别开放词汇目标;(2)基于贝叶斯的价值图将多视角检测置信度融合为每像素的目标相关性分布,通过上限置信区间(UCB)探索实现智能前沿评分;(3)采用去中心化协调策略,结合语义前沿提取、基于测地线路径成本的代价-效用竞标及空间分离惩罚机制,有效减少集群间的冗余探索。

原文摘要 · Abstract (English)

Cooperative visual semantic navigation is a foundational capability for aerial robot teams operating in unknown environments. However, achieving robust open-vocabulary object-goal navigation remains challenging due to the computational constraints of deploying heavy perception models onboard and the complexity of decentralized multi-agent coordination. We present GoalSwarm, a fully decentralized multi-UAV framework for zero-shot semantic object-goal navigation. Each UAV collaboratively constructs a shared, lightweight 2D top-down semantic occupancy map by projecting depth observations from aerial vantage points, eliminating the computational burden of full 3D representations while preserving essential geometric and semantic structure. The core contributions of GoalSwarm are threefold: (1) integration of zero-shot foundation model -- SAM3 for open vocabulary detection and pixel-level segmentation, enabling open-vocabulary target identification without task-specific training; (2) a Bayesian Value Map that fuses multi-viewpoint detection confidences into a per-pixel goal-relevance distribution, enabling informed frontier scoring via Upper Confidence Bound (UCB) exploration; and (3) a decentralized coordination strategy combining semantic frontier extraction, cost-utility bidding with geodesic path costs, and spatial separation penalties to minimize redundant exploration across the swarm.

多无人机语义导航零样本协同决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。