构建城市空间感知多模态数据集,评测AI对城市环境的理解能力。
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

- 基于社交媒体图像构建跨城市多模态数据集,含200万张图。
- 在1K、10K、100K三尺度下验证模型性能变化趋势。
- 支持分类、图文检索、实例分割三类任务,适配城市研究者与AI开发者。
我们提出Urban-ImageNet,一个大规模多模态数据集与评估基准,用于从用户生成的社交媒体图像中分析城市空间感知。该数据集包含2019–2025年间从微博收集的超过200万张公开图像及对应文本,覆盖中国24个城市的61个城区。提供1K、10K、100K三个控制规模的基准子集和完整的200万张全量数据集,适用于大规模训练与评估。数据采用HUSIC框架组织,基于城市理论建立10类分级分类体系,涵盖激活/非激活公共空间、室内外环境、居住、消费内容、人物肖像及非空间社交内容等类别。不同于通用场景数据,本数据集关注机器能否捕捉城市研究中的空间、社会与功能差异。基准支持三项统一任务:(T1)城市场景语义分类,(T2)跨模态图文检索,(T3)实例分割。实验表明,监督分类表现良好,但跨模态检索与实例级分割更具挑战性。多尺度分析揭示模型性能随平衡训练数据从1K增至100K而持续提升。Urban-ImageNet为评估AI对当代城市空间的跨模态理解提供了统一、理论驱动、多城市的基准。数据集与基准代码已公开于huggingface.co/datasets/Yiwei-Ou/Urban-ImageNet 和 github.com/yiasun/dataset-2。
原文摘要 · Abstract (English)
We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imagery. The corpus contains over 2 Million public social media images and paired textual posts collected from Weibo across 61 urban sites in 24 Chinese cities across 2019-2025, with controlled benchmark subsets at 1K, 10K, and 100K scale and a full 2M corpus for large-scale training and evaluation. Urban-ImageNet is organized by HUSIC, a Hierarchical Urban Space Image Classification framework that defines a 10-class taxonomy grounded in urban theory. The taxonomy is designed to distinguish activated and non-activated public spaces, exterior and interior urban environments, accommodation spaces, consumption content, portraits, and non-spatial social-media content. Rather than treating urban imagery as generic scene data, Urban-ImageNet evaluates whether machine perception models can capture spatial, social, and functional distinctions that are central to urban studies. The benchmark supports three tasks within one standardized library: (T1) urban scene semantic classification, (T2) cross-modal image-text retrieval, and (T3) instance segmentation. Our experiments evaluate representative vision, vision-language, and segmentation models, revealing strong performance on supervised scene classification but more challenging behavior in cross-modal retrieval and instance-level urban object segmentation. A multi-scale study further examines how model performance changes as balanced training data increases from 1K, 10K to 100K images. Urban-ImageNet provides a unified, theory-grounded, multi-city benchmark for evaluating how AI systems perceive and interpret contemporary urban spaces across modalities, scales, and task formulations. Dataset and benchmark are available at: huggingface.co/datasets/Yiwei-Ou/Urban-ImageNet and github.com/yiasun/dataset-2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。