通过代表性样本动态优化模型,实现遥感图像分割的零训练渐进提升。
UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution

- 基于代表样本反馈,无须训练即可进化原型
- 在未见类别上仅用少量样本交互即显著提效
- 支持文本与视觉双驱动提示,适配复杂标注流程
提示驱动的视觉语言模型在加速遥感密集标注方面潜力巨大,但静态模型在新场景、未见类别或视觉混淆背景下性能严重退化。现有统一范式多依赖图像内特定提示,缺乏灵活的任务路由机制以适应多意图工作流。实际批量制图中,标注员通常先精修少量代表性样本再处理大规模数据。受此启发,本文提出UniEvo-RS:一种具代表性样本驱动原型演化的全提示遥感分割框架。首先构建统一文本与视觉提示的多指令提示数据集,在单一架构内建立动态任务路由机制,覆盖多样化遥感标注场景。其次引入无训练、反馈驱动的原型进化机制:通过对比人工标注与初始预测在代表性样本上的差异,将预测误差提炼为正负原型,提升大语言模型查询召回率并抑制背景噪声,同时保持固定预算下的聚类记忆。大量实验表明,UniEvo-RS统一多种提示任务,在多数设置下达到当前最优性能。关键在于,仅需对少数样本进行极简交互,即可在批量标注过程中实现对未见类别的零训练渐进精度提升。
原文摘要 · Abstract (English)
Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from severe performance degradation when deployed on novel scenes, unseen categories, or visually confusing backgrounds. Moreover, existing unified paradigms primarily rely on intra-image specific prompts, lacking flexible task routing to adapt to multi-intent operational workflows. In practical batch mapping, annotators typically refine a small set of representative samples before processing large datasets. Motivated by this practice, we propose UniEvo-RS, an omni-prompt unified RS segmentation framework equipped with representative exemplar-driven prototype evolution. First, we construct a multi-instruction prompt dataset that unifies text-driven and visual-driven prompts within a single architecture, establishing a dynamic task-routing mechanism for highly diverse RS annotation scenarios. Second, we introduce a representative feedback-driven, training-free prototype evolution mechanism. By contrasting manual annotations with initial predictions on exemplars, UniEvo-RS distills prediction errors into positive and negative prototypes. These prototypes enhance LLM query recall and suppress spatial background noise under a fixed-budget clustering memory. Extensive experiments show that UniEvo-RS unifies diverse prompting tasks, achieving state-of-the-art performance across most settings. Crucially, with minimal interaction on a few exemplars, it enables training-free, progressive accuracy enhancement on unseen categories during batch annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。