arXiv:2605.26036cs.AIcs.LG2026-05被引 1

构建跨城市多任务统一评估基准,解决城市表征学习中的空间泄露问题。

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

论文配图:CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities
图 1 · 摘自论文原文
  • 采用区块化空间划分避免数据泄露,提升评估可靠性
  • 覆盖8个城市、8类任务,验证模型跨域泛化能力
  • 提供可复现工具链,助力城市基础模型研究

城市表征学习将复杂城市环境编码为通用嵌入,服务于多样化下游任务和新兴的城市基础模型。然而,现有评估方法受限于单一或少数城市与任务,依赖随机划分导致空间泄露,造成性能虚高,削弱跨区域泛化能力与公平比较。为此,我们提出CityRep——一个跨城市、跨任务、跨模态的统一评估基准,采用结构化空间划分策略。该基准包含三部分:(1) 与空间单元无关的评估框架,通过标准化对齐模块支持异构城市表征;(2) 基于区块的空间划分协议,有效缓解空间泄露,实现严谨模型对比;(3) 可扩展的多城市多任务基准套件,涵盖8个城市和8类任务(回归、分类、分布预测)。我们评估了11个代表性城市表征模型。结果表明,模型性能对划分协议高度敏感,随机划分显著抬升分数并改变模型排名。不同城市与任务间表现差异显著,凸显泛化意识评估的重要性。CityRep已开源,包含数据集、评估流程与诊断工具,支持可复现比较,推动城市表征学习向城市基础模型发展。

原文摘要 · Abstract (English)

Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundation models. However, current evaluations are limited, typically focusing on one or two cities and tasks and relying on random splits that introduce spatial leakage, leading to inflated performance and weak support for cross-location generalization and fair comparison. To address this, we propose CityRep, a unified benchmark that evaluates urban representations across data modalities, cities, and tasks using spatially structured splits. CityRep consists of three key components: (1) a spatial unit-agnostic evaluation framework that supports heterogeneous urban representations through a standardized alignment module; (2) a unified evaluation protocol using block-based spatial splits to mitigate spatial leakage and enable rigorous model comparison; and (3) an extensible multi-city, multi-task benchmark suite spanning 8 cities and 8 tasks across regression, classification, and distribution prediction. We evaluate 11 representative urban representation models. Results show that performance is highly sensitive to the split protocol, with random splits inflating scores and altering model rankings. We also observe substantial variability across cities and tasks, underscoring the need for generalization-aware evaluation. CityRep is released as a reproducible benchmark with datasets, evaluation pipelines, and diagnostic tools to facilitate fair comparison and support future research in urban representation learning towards urban foundation models.

城市表征多任务评估空间划分基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。