构建15城24个月的大型轨迹数据集,助力城市移动行为研究。
Massive-STEPS: Massive Semantic Trajectories for Understanding POI Check-ins -- Dataset and Benchmarks
- 基于语义轨迹数据,整合15个跨文化城市的检查点信息
- 覆盖2012-2018年,时长达24个月,超越以往数据集
- 支持监督与零样本模型评估,适合城市规划与智能代理研究
通过兴趣点(POI)轨迹建模理解人类移动行为,在城市规划、个性化服务和生成式智能体模拟中日益重要。然而该领域进展受限于两大挑战:过度依赖2012-2013年的旧数据集,且缺乏可复现、覆盖全球多元城市的市级检查点数据集。为弥补这些空白,我们提出Massive-STEPS(大规模语义轨迹以理解兴趣点打卡),一个基于语义轨迹数据集并增强语义POI元数据的大规模公开基准数据集。Massive-STEPS涵盖15个地理与文化各异的城市,包含2012-2013及2017-2018两个时期的数据,总时长24个月,远超以往数据集。我们在Massive-STEPS上对多种POI模型进行了监督与零样本评估,并在多个城市环境中测试其表现。通过发布Massive-STEPS,我们旨在推动人类移动性与POI轨迹建模的可复现、公平研究。数据集与基准代码已开源:https://github.com/cruiseresearchgroup/Massive-STEPS。
原文摘要 · Abstract (English)
Understanding human mobility through Point-of-Interest (POI) trajectory modeling is increasingly important for applications such as urban planning, personalized services, and generative agent simulation. However, progress in this field is hindered by two key challenges: the over-reliance on older datasets from 2012-2013 and the lack of reproducible, city-level check-in datasets that reflect diverse global regions. To address these gaps, we present Massive-STEPS (Massive Semantic Trajectories for Understanding POI Check-ins), a large-scale, publicly available benchmark dataset built upon the Semantic Trails dataset and enriched with semantic POI metadata. Massive-STEPS spans 15 geographically and culturally diverse cities and features check-in data from both 2012-2013 and 2017-2018, with a longer duration of 24 months than prior datasets. We benchmarked a wide range of POI models on Massive-STEPS using both supervised and zero-shot approaches, and evaluated their performance across multiple urban contexts. By releasing Massive-STEPS, we aim to facilitate reproducible and equitable research in human mobility and POI trajectory modeling. The dataset and benchmarking code are available at: https://github.com/cruiseresearchgroup/Massive-STEPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。