用可解释机器学习方法,将大区域出租车出行数据降尺度到小区域。
Downscaling human mobility data based on demographic socioeconomic and commuting characteristics using interpretable machine learning methods
- 基于人口、经济和通勤特征建模,用四种算法分析出行流与社会因素关系。
- 神经网络在训练集表现最佳,支持向量机在跨区域预测中泛化能力最强。
- 方法可帮助交通规划和城市设计,适合政策制定者和城市研究者使用。
理解城市人类移动模式在不同空间尺度下的特征对社会科学至关重要。本研究提出一种机器学习框架,将纽约市更大空间单元的起点-终点(OD)出租车出行流量降尺度至更小空间单元。首先,利用线性回归(LR)、随机森林(RF)、支持向量机(SVM)和神经网络(NN)四种模型,建立出行流量与人口、社会经济及通勤特征之间的关联。其次,对非线性模型采用扰动敏感性分析以解释变量重要性。结果表明,线性回归模型无法捕捉复杂变量交互;神经网络在训练和测试数据上表现最佳,而支持向量机在降尺度性能上的泛化能力最优。该方法为分析与应用提供了理论进展和实践价值,有助于提升交通服务与城市发展规划。
原文摘要 · Abstract (English)
Understanding urban human mobility patterns at various spatial levels is essential for social science. This study presents a machine learning framework to downscale origin-destination (OD) taxi trips flows in New York City from a larger spatial unit to a smaller spatial unit. First, correlations between OD trips and demographic, socioeconomic, and commuting characteristics are developed using four models: Linear Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and Neural Networks (NN). Second, a perturbation-based sensitivity analysis is applied to interpret variable importance for nonlinear models. The results show that the linear regression model failed to capture the complex variable interactions. While NN performs best with the training and testing datasets, SVM shows the best generalization ability in downscaling performance. The methodology presented in this study provides both analytical advancement and practical applications to improve transportation services and urban development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。