arXiv:2606.15055cs.CVcs.AI2026-06

用分层语义锚点+持续学习,让城市街景模型不再偏爱热门城市。

Bridging Geographic Bias in Urban Streetscape Inference via Lifelong Learning with Visual-Semantic Pivoting

论文配图:Bridging Geographic Bias in Urban Streetscape Inference via Lifelong Learning with Visual-Semantic Pivoting
图 1 · 摘自论文原文
  • 构建三层次语义框架,将街景特征对齐可学习的语义锚点。
  • 在12个城市测试中,跨城市预测相关性达0.834,误差缩小38%。
  • 适合关注城市公平、街景分析与持续学习的研究者使用。

城市街景的视觉感知支撑着景观规划、公共健康与场所营造中的证据决策。然而,基于少数热门大都市训练的模型会系统性误判代表性不足的区域,将地理偏差带入下游政策。我们提出HVSP-LL,一种结合分层视觉-语义锚定模块与公平导向重演机制的持续学习框架。该锚定模块沿宏观结构、中观构成、微观元素三层组织景观概念,并将图像特征对齐各层级的可学习语义锚点,生成抵抗分布漂移的可迁移表征。持续适应组件通过最差区域样本重加权目标与结构感知实例缓冲区,逐步吸收新城市区域并约束跨区域感知差异。我们在涵盖四大洲十二个城市的全景街景基准上评估该框架,其在保留城市序列上的斯皮尔曼相关系数达0.834,较最强持续学习基线绝对提升6.1点,跨城市感知差距降至0.094(相比最强基线减少38%,相比典型正则化基线减少57%)。消融实验表明每一层锚定均贡献单调提升,且公平重演机制使平均反向迁移从-0.038(无保留)转为+0.013,彻底消除保留序列的灾难性遗忘。结果表明,分层锚定是实现城市尺度地理公平街景推断的可行路径。

原文摘要 · Abstract (English)

Visual perception of urban streetscapes underpins evidence-based decisions in landscape planning, public health, and place-making. Yet models trained on a few well-photographed metropolises systematically misjudge underrepresented districts, propagating geographic bias into downstream policy. We address this gap with HVSP-LL, a lifelong learning framework that couples a stratified visual-semantic pivoting module with an equity-aware rehearsal mechanism. The pivoting module organises landscape concepts along a three-tier ontology (macro structure, meso composition, micro element) and aligns image features to learnable semantic anchors at each tier, providing transferable representations that resist distributional drift. The lifelong adaptation component sequentially absorbs new urban regions while constraining inter-region perception gaps through a worst-region sample-reweighting objective and a structurally-aware exemplar buffer. We evaluate HVSP-LL on a panoramic streetscape benchmark assembled from twelve cities across four continents and seven perceptual dimensions. The framework attains 0.834 Spearman correlation on the held-out city sequence, an absolute 6.1 point improvement over the strongest continual baseline, and shrinks the inter-city perception gap to 0.094 -- a 38% reduction relative to the strongest continual baseline (0.151) and a 57% reduction relative to a representative regularisation baseline (0.218). Ablations confirm that each tier of the pivoting hierarchy contributes monotonically, and the equity-aware rehearsal converts mean backward transfer from -0.038 (without retention) to +0.013, eliminating catastrophic forgetting on the held-out sequence. Our results indicate that hierarchical anchoring is a practical pathway toward geographically equitable streetscape inference at city scale.

城市感知持续学习公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。