arXiv:2410.14726cs.LG2024-10

用8年历史数据训练交通预测模型,发现越久远数据越影响精度,提出新方法解决偏差问题。

Incorporating Long-term Data in Training Short-term Traffic Prediction Model

  • 用加权方式对齐数据分布,缓解历史与现状的差异
  • 在真实数据上验证,用96个月数据训练误差显著降低
  • 适用于大多数现有模型,适合长期数据训练场景

短时交通流量预测对智能交通系统至关重要,但多数研究聚焦模型结构优化,忽视了训练数据量的影响。本研究使用纽约八年出租车与共享单车使用数据,评估了最近12、24、48和96个月数据训练的效果。结果发现,使用96个月数据时模型精度反而下降,可能源于协变量偏移(covariate shift)与概念偏移(concept shift)。为此,提出一种基于权重分配的协变量对齐方法,结合环境感知学习应对概念偏移。实验证明,该方法能有效降低测试误差,在大规模历史数据下仍保持高精度。据我们所知,这是首个系统评估连续扩展训练数据对交通预测模型影响的工作。所提方法可嵌入主流短时预测模型,提升其对长期数据的适应能力。

原文摘要 · Abstract (English)

Short-term traffic volume prediction is crucial for intelligent transportation system and there are many researches focusing on this field. However, most of these existing researches concentrated on refining model architecture and ignored amount of training data. Therefore, there remains a noticeable gap in thoroughly exploring the effect of augmented dataset, especially extensive historical data in training. In this research, two datasets containing taxi and bike usage spanning over eight years in New York were used to test such effects. Experiments were conducted to assess the precision of models trained with data in the most recent 12, 24, 48, and 96 months. It was found that the training set encompassing 96 months, at times, resulted in diminished accuracy, which might be owing to disparities between historical traffic patterns and present ones. An analysis was subsequently undertaken to discern potential sources of inconsistent patterns, which may include both covariate shift and concept shift. To address these shifts, we proposed an innovative approach that aligns covariate distributions using a weighting scheme to manage covariate shift, coupled with an environment aware learning method to tackle the concept shift. Experiments based on real word datasets demonstrate the effectiveness of our method which can significantly decrease testing errors and ensure an improvement in accuracy when training with large-scale historical data. As far as we know, this work is the first attempt to assess the impact of contiguously expanding training dataset on the accuracy of traffic prediction models. Besides, our training method is able to be incorporated into most existing short-term traffic prediction models and make them more suitable for long term historical training dataset.

交通预测历史数据数据偏移模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。