arXiv:2607.08004cs.CV2026-07中稿 · SOICT 2025

用文字提示引导无人机图像中定向目标检测,提升复杂场景精度

LOGOS: Language-guided Oriented Object Detection in Aerial Scenes

论文配图:LOGOS: Language-guided Oriented Object Detection in Aerial Scenes
图 1 · 摘自论文原文
  • 通过文本提示动态调节注意力焦点,实现灵活定位
  • 在DOTA数据集上优于现有方法,尤其在密集旋转目标场景
  • 适合遥感图像中需要语义引导的目标检测任务

地理空间场景中的目标检测,如卫星和航拍图像,因目标方向多样、密度不一及背景复杂而面临挑战。传统定向检测方法存在角度不连续、查询尺寸固定及稀疏或杂乱场景效率低下等问题。本文提出LOGOS,一种基于Transformer的新型方法,利用文本提示引导空中场景中的定向目标检测。该方法引入提示调制的内容查询,根据输入文本动态调整模型关注点,从而提升复杂环境下的检测准确率。在DOTA数据集上的大量实验表明,LOGOS显著优于现有最先进方法,尤其在密集排列和旋转目标场景中表现突出。该方法为提升遥感应用中定向目标检测的鲁棒性与可扩展性提供了重要进展。

原文摘要 · Abstract (English)

Object detection in geospatial scenes, such as satellite and aerial imagery, poses significant challenges due to the varying orientations and densities of objects, as well as the complex backgrounds inherent to remote sensing imagery. Traditional methods for oriented object detection have struggled to address issues such as angular discontinuity, fixed query sizes, and inefficiencies in handling sparse or cluttered scenes. In this paper, we propose LOGOS, a novel transformer-based approach that leverages textual prompts to guide the detection of oriented objects in aerial scenes. In particular, our proposed approach incorporates prompt-modulated content queries to dynamically adjust the model's focus based on the provided text, thereby improving object detection accuracy in complex environments. Empirically, extensive experiments on the DOTA dataset demonstrate that LOGOS outperforms existing state-of-the-art methods, particularly in densely packed and rotated object scenarios. Our approach offers a significant step forward in improving the robustness and scalability of oriented object detection in remote sensing applications.

目标检测遥感图像文本引导Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。