CLIP经微调后能更好区分洛杉矶近邻区域,关键在完整场景结构。
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

- 通过微调CLIP编码器提升区域定位能力,尤其依赖完整场景布局。
- 微调后准确率从39%升至82%,预测中心距真实位置缩短至3.86公里。
- 对场景结构敏感,但植被和天空仍主导识别,不靠粗略地理线索。
大规模街景图像包含丰富的城市环境视觉信息,但从中提取精细地理信息仍具挑战,尤其在邻近区域共享粗粒度地理线索时。本文研究洛杉矶八大区域的细粒度区域定位问题,评估预训练CLIP特征在区域区分上的表现及适应后的视觉支持因素。基于9,085张街景图像,对比零样本CLIP、冻结编码器读出、部分编码器更新、低秩适配(LoRA)与全量微调。冻结读出保持约39.03%准确率,而编码器适配达75.94%-82.10%;全微调将预测区域中心均距从12.30公里降至3.86公里。通过语义线索移除、边缘图与模糊处理降低外观、以及区块打乱破坏场景配置进行探测。适配模型在边缘和模糊任务中表现更优,且在打乱后42.92%-45.56%预测发生变化,远高于冻结方法的10.79%-14.60%。然而,外观减弱下性能保留比例未提升,植被与天空仍具影响力。Caltech101对照实验表明,场景打乱敏感性非地理定位特有。总体而言,编码器适配显著提升邻近区域判别力,关联于对完整场景配置的更强敏感性,而非仅依赖粗略结构。结论适用于已知地点附近的视角变化,而非地理上分离的泛化。
原文摘要 · Abstract (English)
Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination, and what visual information supports performance after adaptation. Using 9,085 street-view images from eight Greater Los Angeles regions, we compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Frozen readouts remain near the 39.03% zero-shot accuracy, whereas encoder adaptation achieves 75.94-82.10%. Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km. We probe these gains through semantic cue removal, appearance reduction using edge maps and blur, and scene-configuration disruption using patch scrambling. Adapted models achieve higher edge and blur accuracy and switch 42.92-45.56% of predictions after scrambling, compared with 10.79-14.60% for frozen methods. However, adaptation does not improve the fraction of performance retained after appearance reduction, while vegetation and sky remain influential. A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization. Overall, encoder adaptation substantially improves nearby-region discrimination and is associated with greater sensitivity to intact scene configuration, without evidence that coarse structure alone becomes sufficient for prediction. These conclusions concern viewpoint variation near known locations rather than geographically disjoint generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。