用加权损失提升交通数据补全精度,更抗异常值干扰。
Latent Factorization of Tensors with Threshold Distance Weighted Loss for Traffic Data Estimation
- 引入阈值距离加权损失,对不同样本分配差异化权重
- 在两个城市数据集上预测误差显著低于现有方法
- 适合处理含异常值的实时交通数据补全任务
智能交通系统(ITS)依赖完整高质量的时空交通数据以实现最优性能。然而,现实中的通信故障和传感器失灵常导致数据缺失或损坏,给ITS发展带来挑战。在多种时空数据补全方法中,张量潜在因子分解(LFT)模型被广泛采用且效果良好。但传统LFT模型通常使用标准L2范数作为损失函数,易受异常值影响。为此,本文提出一种融入阈值距离加权(TDW)损失的张量潜在因子分解(TDWLFT)模型。该损失函数通过为不同样本分配差异化权重,有效降低模型对异常值的敏感性。在两个来自不同城市环境的交通速度数据集上进行的大量实验表明,所提TDWLFT模型在预测精度和计算效率方面均持续优于当前先进方法。
原文摘要 · Abstract (English)
Intelligent transportation systems (ITS) rely heavily on complete and high-quality spatiotemporal traffic data to achieve optimal performance. Nevertheless, in real-word traffic data collection processes, issues such as communication failures and sensor malfunctions often lead to incomplete or corrupted datasets, thereby posing significant challenges to the advancement of ITS. Among various methods for imputing missing spatiotemporal traffic data, the latent factorization of tensors (LFT) model has emerged as a widely adopted and effective solution. However, conventional LFT models typically employ the standard L2-norm in their learning objective, which makes them vulnerable to the influence of outliers. To overcome this limitation, this paper proposes a threshold distance weighted (TDW) loss-incorporated Latent Factorization of Tensors (TDWLFT) model. The proposed loss function effectively reduces the model's sensitivity to outliers by assigning differentiated weights to individual samples. Extensive experiments conducted on two traffic speed datasets sourced from diverse urban environments confirm that the proposed TDWLFT model consistently outperforms state-of-the-art approaches in terms of both in both prediction accuracy and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。