arXiv:2510.17670cs.LGcs.AI2025-10被引 1

用少量标注快速适配遥感图像中的开放词汇检测,提升细粒度识别准确率。

On-the-Fly OVD Adaptation with FLAME: Few-shot Localization via Active Marginal-Samples Exploration

  • 通过主动学习选取关键样本,实时训练轻量分类器
  • 仅需数个标注样本,1分钟内完成适应,精度超越现有方法
  • 适合遥感等专业领域用户快速定制检测需求

开放词汇目标检测(OVD)模型可通过任意文本查询检测物体,但其在遥感(RS)等专业领域的零样本性能常受自然语言模糊性影响。例如,'渔船'与'游艇'的语义嵌入相似,难以区分,阻碍非法捕捞监测等应用。为此,我们提出一种级联方法:先用预训练零样本模型生成高召回候选框,再通过实时训练的轻量分类器进行高精度修正,仅需少量用户标注,大幅降低遥感图像标注成本。核心为FLAME——一种一步式主动学习策略,利用密度估计识别决策边界附近的不确定边缘样本,并通过聚类保证多样性。该采样方法无需全模型微调,可在1分钟内实现即时适应,显著快于当前最优方案。在多个遥感基准上持续优于现有方法,构建了高效实用的通用模型定制框架。

原文摘要 · Abstract (English)

Open-vocabulary object detection (OVD) models offer remarkable flexibility by detecting objects from arbitrary text queries. However, their zero-shot performance in specialized domains like Remote Sensing (RS) is often compromised by the inherent ambiguity of natural language, limiting critical downstream applications. For instance, an OVD model may struggle to distinguish between fine-grained classes such as "fishing boat" and "yacht" since their embeddings are similar and often inseparable. This can hamper specific user goals, such as monitoring illegal fishing, by producing irrelevant detections. To address this, we propose a cascaded approach that couples the broad generalization of a large pre-trained OVD model with a lightweight few-shot classifier. Our method first employs the zero-shot model to generate high-recall object proposals. These proposals are then refined for high precision by a compact classifier trained in real-time on only a handful of user-annotated examples - drastically reducing the high costs of RS imagery annotation.The core of our framework is FLAME, a one-step active learning strategy that selects the most informative samples for training. FLAME identifies, on the fly, uncertain marginal candidates near the decision boundary using density estimation, followed by clustering to ensure sample diversity. This efficient sampling technique achieves high accuracy without costly full-model fine-tuning and enables instant adaptation, within less then a minute, which is significantly faster than state-of-the-art alternatives.Our method consistently surpasses state-of-the-art performance on RS benchmarks, establishing a practical and resource-efficient framework for adapting foundation models to specific user needs.

开放词汇检测遥感图像主动学习少样本适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。