用少量样本生成更精准的遥感图像文本查询,提升开放词汇分割效果
Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion

- 通过文本反转技术从少量样本中提取更有效的查询词
- 在基准测试上将平均交并比从3.9提升至39.4
- 适合需要快速适配新类别的遥感图像分割场景
开放词汇分割无需每类训练即可通过文本查询识别任意类别,但在遥感图像上表现不佳。我们发现问题主要出在文本查询本身——因模型未针对航拍图像优化,普通类别名作为查询在视觉-语言嵌入空间中定位不准。改进查询名称可缓解部分性能差距,但现有自然语言改写无法解决剩余问题。本文通过在冻结模型上对少量样本进行文本反转,仅保留推理时的文本查询,重建有效语义地址。在代表性基准测试中,受影响类别的平均交并比从3.9提升至39.4;在八个遥感数据集上,优于依赖视觉提示的少样本方法。
原文摘要 · Abstract (English)
Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on categories it handles reliably elsewhere. We find that much of this gap traces to the text query rather than to the segmentation model. Because these models are not specialized for overhead imagery, the class name that serves as the query is often a weak address into the vision-language embedding space. We show that a better name repairs part of the gap, while the remaining failures call for an address that the tested natural-language rephrasings do not provide. We recover that address from a few examples through textual inversion on a frozen model, keeping inference text only. On a representative benchmark this raises the mean intersection over union on the affected categories from 3.9 to 39.4, and across eight remote sensing datasets it improves over few-shot methods that instead inject visual prompts at inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。