用噪声抑制与长文本对齐提升城市图像表征学习效果
Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
- 通过插值对齐长文本与城市图像细粒度特征
- 在四座真实城市上实现优于基线的跨模态性能
- 适合从事城市计算、多模态学习的研究者
区域表征学习在城市计算中至关重要,能从无标签城市数据中提取有意义特征。如同面部年龄反映个体健康状况,城市的视觉外观构成其“肖像”,蕴含潜在的社会经济与环境特征。近期研究尝试利用大语言模型(LLMs)将文本知识融入基于图像的城市区域表征学习。然而仍面临两大挑战:一、细粒度视觉特征与长文本难以对齐;二、因LLM生成文本存在噪声,导致知识融合效果不佳。为此,我们提出一种名为UrbanLN的新预训练框架,通过长文本感知与噪声抑制改进城市区域表征学习。具体地,引入信息保持的拉伸插值策略,实现复杂城市场景中长文本与细粒度视觉语义的对齐。为有效挖掘LLM生成文本中的知识并过滤噪声,提出双层优化策略:数据层面,构建多模型协作流水线,自动产生多样且可靠的文本描述,无需人工干预;模型层面,采用基于动量的自蒸馏机制生成稳定伪目标,支持在噪声条件下稳健的跨模态学习。在四个真实城市及多种下游任务上的大量实验表明,UrbanLN性能显著优于基线。
原文摘要 · Abstract (English)
Region representation learning plays a pivotal role in urban computing by extracting meaningful features from unlabeled urban data. Analogous to how perceived facial age reflects an individual's health, the visual appearance of a city serves as its "portrait", encapsulating latent socio-economic and environmental characteristics. Recent studies have explored leveraging Large Language Models (LLMs) to incorporate textual knowledge into imagery-based urban region representation learning. However, two major challenges remain: i) difficulty in aligning fine-grained visual features with long captions, and ii) suboptimal knowledge incorporation due to noise in LLM-generated captions. To address these issues, we propose a novel pre-training framework called UrbanLN that improves Urban region representation learning through Long-text awareness and Noise suppression. Specifically, we introduce an information-preserved stretching interpolation strategy that aligns long captions with fine-grained visual semantics in complex urban scenes. To effectively mine knowledge from LLM-generated captions and filter out noise, we propose a dual-level optimization strategy. At the data level, a multi-model collaboration pipeline automatically generates diverse and reliable captions without human intervention. At the model level, we employ a momentum-based self-distillation mechanism to generate stable pseudo-targets, facilitating robust cross-modal learning under noisy conditions. Extensive experiments across four real-world cities and various downstream tasks demonstrate the superior performance of our UrbanLN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。