用提示词提升小规模光伏分割精度,仅需少量标注即可高效部署。
Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing Imagery

- 通过文本、几何和混合提示对比实验,发现空间提示更有效。
- 使用几百个标注样本即实现高精度,数据效率显著。
- 适合在标注稀缺的偏远地区进行大规模光伏设施测绘。
时空光伏数据对理解离网地区光伏采用过程至关重要,但现有数据仍严重缺乏。遥感图像自动分割提供了可行方案,但居民区光伏系统因尺寸小、分布稀疏,导致目标与背景严重失衡。视觉-语言基础模型(FMs)通过基于提示的语义与空间引导实现数据高效建模,但不同提示类型的作用尚不明确。本研究系统评估了SAM3在遥感图像中对小规模光伏分割的表现,比较了文本、几何及混合提示,在不同监督水平、训练策略、空间分辨率和成像条件下的表现。以大型离网农村区域的多时相航拍影像为研究区,并在三个额外数据集上验证结果。提示策略是决定模型行为的主导因素:文本提示性能最低,且对监督程度和成像条件敏感;空间引导显著提升分割精度与鲁棒性;混合提示达到最高准确率与稳定性,表明语义与空间信息具有互补性。多数性能提升仅需数百个标注样本,体现强数据效率。迁移学习整体影响有限,仅在低监督下对文本提示有小幅改善。总体而言,提示策略是决定SAM3适应性、鲁棒性与泛化能力的关键,凸显可提示基础模型在数据受限离网地区规模化光伏制图中的潜力。
原文摘要 · Abstract (English)
Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segmentation of remote sensing (RS) imagery offers a promising solution; yet, residential PV systems remain challenging targets because of their small size and sparse distribution, resulting in severe target-background imbalance. Vision-language foundation models (FMs) provide a data-efficient paradigm through prompt-based semantic and spatial guidance, but the relative contribution of different prompt types remains unclear. We systematically evaluate SAM3 for small-scale PV segmentation in RS imagery by comparing textual, geometric, and hybrid prompting, under varying supervision levels, training strategies, spatial resolutions, and imaging conditions. Multi-temporal aerial imagery from a large off-grid rural region serves as a study site, with findings validated across three additional datasets. Prompting strategy emerged as the dominant factor governing model behavior. Textual prompting consistently produced the lowest performance and showed the greatest sensitivity to supervision and imaging conditions. In contrast, spatial guidance substantially improved both segmentation accuracy and robustness. Hybrid prompting achieved the highest accuracy and stability, indicating that semantic and spatial guidance provide complementary information. Most performance gains were achieved with only a few hundred annotated samples, demonstrating strong data efficiency. Transfer learning had limited overall impact, with only modest improvements observed for textual prompting under limited supervision. Overall, our findings establish prompting strategy as a key determinant of SAM3 adaptation, robustness, and generalization, highlighting the potential of promptable FMs for scalable PV mapping in data-constrained off-grid regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。