用少量标注数据微调,比提示工程更有效于卫星云分割。
Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift

- 仅需0.1%标签数据(约8张图)微调,就超越零样本表现。
- 0.5%-1%数据时出现性能下降,但最终恢复并超过提示方法。
- 适合想在稀缺标注下部署视觉语言模型的遥感研究者。
将视觉语言模型适配至遥感影像面临根本挑战:卫星数据的视觉与语言分布远偏离自然图像预训练语料。尽管如此,提示仍是主流部署范式,源于领域特定语言可引导冻结模型表征的假设。我们在云分割任务上直接检验该假设,使用CLIPSeg在CloudSEN12+基准测试60种提示变体,涵盖简单标签、领域术语、外观描述和上下文线索。结果发现,所有提示均低于零样本基线(0.255 mIoU),优化提示甚至低至0.07 mIoU。语言细化无法弥合自然图像表征与卫星光谱影像间的差距。相反,仅用0.1%标签数据(约8张图)的监督微调整体优于零样本,5%-10%数据即可恢复约85%最大可达成mIoU。全量微调始终优于低秩适配,差距达0.03-0.09 mIoU,尤其在光谱模糊类别中显著。在0.5%-1%数据时,这些类别先出现性能下降,随后恢复,而整体mIoU可能掩盖此现象。对遥感领域适配视觉语言模型的实践者而言,结论明确:标注数据不是提示的替代品,而是值得投入的路径。
原文摘要 · Abstract (English)
Adapting vision-language models to remote sensing imagery presents a fundamental challenge: both the visual and linguistic distributions of satellite data lie far outside natural image pretraining corpora. Despite this, prompting remains the dominant deployment paradigm, driven by the assumption that domain-specific language can guide frozen model representations toward specialized tasks. We test this assumption directly on a domain where the mismatch is prominent: cloud segmentation for satellite imagery. Using CLIPSeg on the CloudSEN12+ cloud segmentation benchmark, we evaluate 60 prompt variants spanning simple labels, domain terminology, appearance descriptors, and contextual cues, finding that every variant underperforms the zero-shot baseline (0.255 mIoU), with engineered prompts scoring as low as 0.07 mIoU. No amount of linguistic refinement bridges the gap between CLIP's natural image representations and satellite spectral imagery. In contrast, supervised fine-tuning with just 0.1% labeled data (~8 images) surpasses zero-shot performance overall, and 5-10% data recovers ~85% of maximum achievable mIoU. Full fine-tuning consistently outperforms low-rank adaptation by 0.03-0.09 mIoU, with the largest gaps for spectrally ambiguous classes, and at 0.5 to 1% labeled data, fine-tuning temporarily degrades performance on these classes before recovering, a supervision dip that aggregate mIoU can mask. For practitioners adapting vision-language models to specialized imagery, our results deliver a clear message: labeled data is not the expensive alternative to prompting; it is the worthwhile path.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。