arXiv:2502.00392cs.CV2025-02被引 18

构建首个无人机场景指代理解基准,解决小目标与复杂环境难题。

RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes

  • 用多智能体框架半自动标注,提升标注效率与表达质量。
  • 提出NGDINO模型,显式学习语言中物体数量,提升多目标与无目标识别性能。
  • 适合研究无人机视觉语言、多目标检测与场景理解的学者使用。

无人机作为广泛应用的机器人平台,在具身人工智能中展现出巨大潜力。指代表达理解(REC)使无人机能根据自然语言定位物体,是具身AI的关键能力。尽管地面场景的REC已取得进展,但航拍视角带来视角变化、遮挡和尺度差异等独特挑战。为此,我们提出RefDrone——首个面向无人机场景的REC基准数据集。该数据集揭示三大核心挑战:1)多尺度及小目标检测;2)多目标与无目标样本共存;3)环境复杂且依赖丰富上下文表达。为高效构建数据集,我们开发了RDAgent(基于多智能体系统的指代无人机标注框架),实现高质量上下文表达生成并降低标注成本。此外,提出Number Grounding DINO(NGDINO)方法,显式建模语言中提及的物体数量。在多种前沿REC模型上的全面实验表明,NGDINO在新提出的RefDrone和现有gRefCOCO数据集上均表现更优。数据集与代码已公开于https://github.com/sunzc-sunny/refdrone。

原文摘要 · Abstract (English)

Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comprehension (REC) enables drones to locate objects based on natural language expressions, a crucial capability for Embodied AI. Despite advances in REC for ground-level scenes, aerial views introduce unique challenges including varying viewpoints, occlusions and scale variations. To address this gap, we introduce RefDrone, a REC benchmark for drone scenes. RefDrone reveals three key challenges in REC: 1) multi-scale and small-scale target detection; 2) multi-target and no-target samples; 3) complex environment with rich contextual expressions. To efficiently construct this dataset, we develop RDAgent (referring drone annotation framework with multi-agent system), a semi-automated annotation tool for REC tasks. RDAgent ensures high-quality contextual expressions and reduces annotation cost. Furthermore, we propose Number GroundingDINO (NGDINO), a novel method designed to handle multi-target and no-target cases. NGDINO explicitly learns and utilizes the number of objects referred to in the expression. Comprehensive experiments with state-of-the-art REC methods demonstrate that NGDINO achieves superior performance on both the proposed RefDrone and the existing gRefCOCO datasets. The dataset and code are be publicly at https://github.com/sunzc-sunny/refdrone.

无人机指代理解多目标检测视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。