arXiv:2601.00998cs.CV2026-01被引 5

构建无人机影像隐式视觉定位基准,提升模型推理能力

DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models

  • 提出隐式转显式思维链,结合强化学习增强场景理解
  • 涵盖6类场景,每对象含显式与隐式查询,共3.2万样本
  • 揭示主流模型在隐式指代上严重不足,适合无人机智能研究者

遥感大视觉语言模型(LVLM)在视觉定位(VG)任务中展现出强大潜力。然而,现有遥感视觉定位数据集主要依赖显式指代表达,如相对位置、大小和颜色线索,限制了对需场景知识的隐式定位任务的表现。本文提出DVGBench,一个高质量的无人机影像隐式视觉定位基准,覆盖交通、灾害、安防、体育、社交活动和生产活动六类主要应用场景。每个目标同时提供显式与隐式查询,共包含3.2万条标注样本。基于该数据集,设计了DroneVG-R1模型,其在强化学习框架下集成新型隐式转显式思维链(I2E-CoT),利用场景特定知识将隐式引用转化为显式描述,从而降低定位难度。对主流模型在显式与隐式任务上的评估显示,其推理能力存在显著不足。这些发现为提升无人机智能体的推理能力提供了切实可行的指导。代码与数据集将在https://github.com/zytx121/DVGBench发布。

原文摘要 · Abstract (English)

Remote sensing (RS) large vision-language models (LVLMs) have shown strong promise across visual grounding (VG) tasks. However, existing RS VG datasets predominantly rely on explicit referring expressions-such as relative position, relative size, and color cues-thereby constraining performance on implicit VG tasks that require scenario-specific domain knowledge. This article introduces DVGBench, a high-quality implicit VG benchmark for drones, covering six major application scenarios: traffic, disaster, security, sport, social activity, and productive activity. Each object provides both explicit and implicit queries. Based on the dataset, we design DroneVG-R1, an LVLM that integrates the novel Implicit-to-Explicit Chain-of-Thought (I2E-CoT) within a reinforcement learning paradigm. This enables the model to take advantage of scene-specific expertise, converting implicit references into explicit ones and thus reducing grounding difficulty. Finally, an evaluation of mainstream models on both explicit and implicit VG tasks reveals substantial limitations in their reasoning capabilities. These findings provide actionable insights for advancing the reasoning capacity of LVLMs for drone-based agents. The code and datasets will be released at https://github.com/zytx121/DVGBench

视觉定位无人机大模型隐式指代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。