arXiv:2409.16837cs.LGcs.CY2024-09被引 4

用收入等基础人口数据提升城市区域嵌入质量,预测更准。

Demo2Vec: Learning Region Embedding with Demographic Information

  • 结合收入等人口数据,改进区域嵌入的预训练方法
  • 纽约和芝加哥实验显示性能最高提升10.22%
  • 适合缺乏移动大数据的发展中城市使用

人口数据(如收入、教育水平、就业率)蕴含丰富的城市区域信息,但鲜有研究将其融入区域嵌入生成。本研究证明,简单易得的人口数据可显著提升先进区域嵌入模型的质量,并在三种常见城市任务——签到预测、犯罪率预测和房价预测中取得更优表现。我们发现,基于KL散度的现有预训练方法可能偏向于移动性信息,因此提出采用Jensen-Shannon散度作为多视角表示学习更合适的损失函数。纽约和芝加哥的实验表明,移动性+收入是最佳预训练数据组合,相比现有模型预测性能最高提升10.22%。考虑到许多发展中国家难以获取移动大数据,建议采用地理邻近性+收入这一简单但有效的数据组合进行区域嵌入预训练。

原文摘要 · Abstract (English)

Demographic data, such as income, education level, and employment rate, contain valuable information of urban regions, yet few studies have integrated demographic information to generate region embedding. In this study, we show how the simple and easy-to-access demographic data can improve the quality of state-of-the-art region embedding and provide better predictive performances in urban areas across three common urban tasks, namely check-in prediction, crime rate prediction, and house price prediction. We find that existing pre-train methods based on KL divergence are potentially biased towards mobility information and propose to use Jenson-Shannon divergence as a more appropriate loss function for multi-view representation learning. Experimental results from both New York and Chicago show that mobility + income is the best pre-train data combination, providing up to 10.22\% better predictive performances than existing models. Considering that mobility big data can be hardly accessible in many developing cities, we suggest geographic proximity + income to be a simple but effective data combination for region embedding pre-training.

区域嵌入人口数据城市预测多视图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。