arXiv:2511.21053cs.ROcs.CV2025-11AAAI被引 8

首个无人机场景下基于语言指令的多目标跟踪基准与方法

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

  • 构建无人机视角下的语言引导多目标跟踪数据集
  • 提出协同视觉-语言表征学习模型,提升空中场景理解能力
  • 适合研究空地协同智能、视觉语言模型与无人机应用的学者

指代式多目标跟踪(RMOT)旨在通过自然语言指令实现精准的目标检测与追踪,是智能机器人系统的核心能力。然而,当前研究主要集中于地面场景,难以捕捉大范围场景上下文,限制了全面追踪与路径规划能力。相比之下,无人机(UAV)凭借广阔的空中视野和优越的机动性,可实现大范围监控。此外,无人机已成为具身智能的关键平台,对具备自然语言交互能力的智能空中系统提出了前所未有的需求。为此,我们提出AerialMind,首个面向无人机场景的大规模RMOT基准数据集,旨在填补这一研究空白。为支持其构建,我们开发了创新的半自动化协作式代理标注框架(COALA),显著降低人工成本并保持标注质量。同时,我们提出HawkEyeTrack(HETrack)方法,通过协同增强视觉-语言表征学习,提升对无人机场景的感知能力。大量实验验证了数据集的挑战性及方法的有效性。

原文摘要 · Abstract (English)

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrains their ability to capture broad-scale scene contexts and perform comprehensive tracking and path planning. In contrast, Unmanned Aerial Vehicles (UAVs) leverage their expansive aerial perspectives and superior maneuverability to enable wide-area surveillance. Moreover, UAVs have emerged as critical platforms for Embodied Intelligence, which has given rise to an unprecedented demand for intelligent aerial systems capable of natural language interaction. To this end, we introduce AerialMind, the first large-scale RMOT benchmark in UAV scenarios, which aims to bridge this research gap. To facilitate its construction, we develop an innovative semi-automated collaborative agent-based labeling assistant (COALA) framework that significantly reduces labor costs while maintaining annotation quality. Furthermore, we propose HawkEyeTrack (HETrack), a novel method that collaboratively enhances vision-language representation learning and improves the perception of UAV scenarios. Comprehensive experiments validated the challenging nature of our dataset and the effectiveness of our method.

多目标跟踪无人机视觉语言具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。