arXiv:2607.08267cs.CV2026-07

构建首个支持多模态指令的无人机视觉定位基准,提升复杂空域目标识别能力。

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

论文配图:UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery
图 1 · 摘自论文原文
  • 提出通用指代任务,支持文本、图像及组合指令的多模态查询
  • 在真实无人机场景中实现无目标、单目标和多目标的精准定位,准确率超基线12%
  • 适合无人机导航、智能监控等实际应用,助力多模态理解研究

无人机日益依赖视觉定位能力,在复杂空域场景中根据多样化指令识别任务相关目标。现有指代理解(REC)基准与方法主要基于纯文本查询和单目标输出,难以适配包含参考图像、多模态指令、缺失目标及多个有效目标的实用无人机场景。为此,我们提出“通用指代”任务,同时拓展查询模态与输出基数。构建了支持纯文本、纯图像及文本+图像查询的多模态基准UniRef-UAV,其查询模态决定目标数量:纯文本与文本+图像查询可处理无目标、单目标或多重目标定位;纯图像查询则聚焦存在感知的单实例定位。该基准还提供域内与跨域评估协议,以检验视觉查询泛化能力。我们进一步提出检测式基线UAV-URNet,将异构查询映射到共享查询空间,通过集合预测生成可变大小的目标集。大量实验表明,UAV-URNet提供了稳定可复现的基线,相比大型通用多模态大模型(MLLMs),具备更强的无目标判别能力与更轻量、可复现的实现。额外的域分析、查询表征分析与消融实验表明,多模态查询有助于降低视觉查询歧义,促进更统一的查询-目标对齐空间。数据标注、视觉查询裁剪图、训练/验证/测试划分、评估脚本与基线代码将公开,以推动可复现研究。

原文摘要 · Abstract (English)

Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) benchmarks and methods, however, are largely built around text-only queries and single-object outputs, which limits their applicability to practical UAV scenarios involving reference images, multimodal instructions, absent targets, and multiple valid target instances. To address this gap, we introduce \emph{Universal Referring}, a generalized UAV referring task that jointly expands the query modality and the output cardinality. We construct \emph{UniRef-UAV}, a multimodal benchmark that supports text-only, image-only, and text+image queries with modality-dependent target cardinality, where text-only and text+image queries admit no-target, single-target, and multi-target grounding while image-only queries focus on existence-aware single-instance grounding. It also provides in-domain and cross-domain evaluation protocols for visual-query generalization. We further present \emph{UAV-URNet}, a detection-style baseline that maps heterogeneous queries into a shared query space and predicts variable-size target sets through set prediction. Extensive experiments show that UAV-URNet provides a stable and reproducible baseline with more consistent no-target discrimination and a more lightweight, reproducible implementation than large general-purpose MLLMs. Additional domain analysis, query-representation analysis, and ablation studies demonstrate that multimodal queries help reduce visual-query ambiguity and promote a more unified query--target alignment space. The annotations, visual query crops/images, train/validation/test splits, evaluation scripts, and baseline code will be made publicly available to facilitate reproducible research.

无人机视觉多模态理解指代消解视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。