arXiv:2503.06146cs.CV2025-03ICCV被引 12

提出开放提示遥感目标检测框架,支持多模态提示与实时高精度检测。

OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images

  • 引入多模态提示与多任务检测头,平衡精度与速度
  • 在7个数据集上平均精度比YOLO-World高8.7%,达20.8 FPS
  • 适合需要实时、泛化强的遥感图像分析场景

遥感目标检测虽取得显著进展,但多数研究仍聚焦于封闭集检测,限制了跨数据集的泛化能力。开放词汇目标检测(OVD)通过文本提示与视觉特征的多模态关联提供解决方案。然而,现有遥感图像的OVD方法受限于小规模数据集,且未能应对遥感解读的独特挑战,如定向目标检测,以及在多样化场景下对高精度与实时性的双重需求。为此,我们提出OpenRSD——一种通用的开放提示遥感目标检测框架。该框架支持多模态提示,并集成多任务检测头以兼顾准确率与实时性。此外,设计了多阶段训练流程以增强模型泛化能力。在七个公开数据集上的评估表明,OpenRSD在定向与水平边界框检测中均表现优异,具备实时推理能力,适用于大规模遥感图像分析。相比YOLO-World,其平均精度提升8.7%,推理速度达20.8 FPS。代码与模型将公开。

原文摘要 · Abstract (English)

Remote sensing object detection has made significant progress, but most studies still focus on closed-set detection, limiting generalization across diverse datasets. Open-vocabulary object detection (OVD) provides a solution by leveraging multimodal associations between text prompts and visual features. However, existing OVD methods for remote sensing (RS) images are constrained by small-scale datasets and fail to address the unique challenges of remote sensing interpretation, include oriented object detection and the need for both high precision and real-time performance in diverse scenarios. To tackle these challenges, we propose OpenRSD, a universal open-prompt RS object detection framework. OpenRSD supports multimodal prompts and integrates multi-task detection heads to balance accuracy and real-time requirements. Additionally, we design a multi-stage training pipeline to enhance the generalization of model. Evaluated on seven public datasets, OpenRSD demonstrates superior performance in oriented and horizontal bounding box detection, with real-time inference capabilities suitable for large-scale RS image analysis. Compared to YOLO-World, OpenRSD exhibits an 8.7\% higher average precision and achieves an inference speed of 20.8 FPS. Codes and models will be released.

遥感检测开放词汇多模态实时检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。