arXiv:2609.05953cs.CV2026-09

用原型记忆增强大模型,提升遥感图像少样本细粒度目标检测性能

ProtoRAG: Prototype-Based Retrieval Augmentation for Few-Shot Fine-Grained Remote Sensing Object Detection

论文配图:ProtoRAG: Prototype-Based Retrieval Augmentation for Few-Shot Fine-Grained Remote Sensing Object Detection
图 1 · 摘自论文原文
  • 通过原型记忆库分离定位与细粒度识别任务
  • 在三个数据集上最高提升14.80 mAP50,显著优于基线
  • 适合少样本、细粒度遥感目标检测场景

遥感图像中的少样本细粒度目标检测面临标注稀缺的挑战,需同时实现精确定位与视觉相似子类的区分。尽管多模态大模型能提供粗略定位,但缺乏可靠的细粒度视觉证据。为此,我们提出ProtoRAG,一种基于原型的检索增强框架,通过外部物体级视觉记忆将粗略定位与细粒度识别解耦。为从有限支持样本构建可靠视觉记忆,引入判别性原型空间学习(DPSL),通过监督对比学习和原型一致性正则化,促进判别性强且稳定的表征。进一步设计不确定性引导的候选约束推理策略,仅对模糊实例检索特定候选视觉参考并触发多模态推理。大量实验表明,ProtoRAG在九种少样本设置下持续超越代表性基线,在MAR20、HRSC2016和FAIR1M-2.0上分别领先最强基线14.80、2.27和4.04 mAP$_{50}$。

原文摘要 · Abstract (English)

Few-shot fine-grained object detection (FGOD) in remote sensing imagery is challenging because limited annotations must support both object localization and discrimination among visually similar subcategories. Although multimodal large language models (MLLMs) provide strong coarse object localization, they lack explicit visual evidence for reliable fine-grained recognition. To address this limitation, we propose ProtoRAG, a prototype-based retrieval-augmented framework that decouples coarse localization from fine-grained recognition by equipping MLLMs with an external object-level visual memory. To construct a reliable visual memory from limited support samples, we introduce Discriminative Prototype Space Learning (DPSL), which encourages discriminative and prototype-stable representations through supervised contrastive learning and prototype-consistency regularization. We further develop an uncertainty-guided candidate-constrained reasoning strategy that augments MLLMs with retrieved candidate-specific visual references and invokes multimodal reasoning only for ambiguous instances. Extensive experiments show that ProtoRAG consistently surpasses representative baselines in nine few-shot settings, outperforming the strongest baselines by 14.80, 2.27, and 4.04 mAP$_{50}$ on MAR20, HRSC2016, and FAIR1M-2.0, respectively.

遥感检测少样本学习多模态原型记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。