arXiv:2512.07302cs.CVcs.AI2025-12被引 1

用智能提示增强提升无人机图像理解准确率

Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts

  • 设计Agent框架自动提取多维辅助信息优化任务提示
  • 在复杂无人机图像上实现开源自研模型性能显著提升
  • 适合需要高精度无人机视觉理解的研究与应用者

基于视觉语言模型(VLM)的图像感知方法通常依赖用户提供的文本任务提示来提取和分析图像内容。然而,该方法在无人机影像中表现受限,主要因目标混淆、尺度变化和复杂背景导致视觉与文本标记语义对齐困难。为解决此问题,本文提出AerialVP——首个面向无人机图像感知的任务提示增强代理框架。AerialVP主动从无人机图像中提取多维辅助信息以增强任务提示,包含三个阶段:分析任务提示识别任务类型与增强需求、从工具库中选择合适工具、结合分析结果与工具生成增强提示。为评估AerialVP,构建AerialSense基准测试集,涵盖航空视觉推理、问答与定位任务,支持在多种分辨率、光照条件及城市/自然场景下的通用性评估。实验表明,AerialVP显著提升提示引导能力,在开源与专有VLM上均实现稳定且显著的性能改进。

原文摘要 · Abstract (English)

Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, such methods face limitations when applied to UAV imagery, which presents challenges like target confusion, scale variations, and complex backgrounds. These challenges arise because VLMs' understanding of image content depends on the semantic alignment between visual and textual tokens. When the task prompt is simplistic and the image content is complex, achieving effective alignment becomes difficult, limiting the model's ability to focus on task-relevant information. To address this issue, we introduce AerialVP, the first agent framework for task prompt enhancement in UAV image perception. AerialVP proactively extracts multi-dimensional auxiliary information from UAV images to enhance task prompts, overcoming the limitations of traditional VLM-based approaches. Specifically, the enhancement process includes three stages: (1) analyzing the task prompt to identify the task type and enhancement needs, (2) selecting appropriate tools from the tool repository, and (3) generating enhanced task prompts based on the analysis and selected tools. To evaluate AerialVP, we introduce AerialSense, a comprehensive benchmark for UAV image perception that includes Aerial Visual Reasoning, Aerial Visual Question Answering, and Aerial Visual Grounding tasks. AerialSense provides a standardized basis for evaluating model generalization and performance across diverse resolutions, lighting conditions, and both urban and natural scenes. Experimental results demonstrate that AerialVP significantly enhances task prompt guidance, leading to stable and substantial performance improvements in both open-source and proprietary VLMs. Our work will be available at https://github.com/lostwolves/AerialVP.

无人机视觉提示工程视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。