arXiv:2410.04740cs.LGcs.AI2024-10被引 14

提出时空模型泛化能力评估新基准,发现主流模型在真实城市环境变化下表现大幅下降。

Evaluating the Generalization Ability of Spatiotemporal Model in Urban Scenario

  • 构建六类城市场景的分布外测试集,涵盖跨年数据变化
  • 多数模型在跨年数据上性能远低于简单MLP,显著退化
  • 适度丢弃率可有效提升泛化能力,兼顾分布内表现仍具挑战

时空神经网络在城市场景中展现出捕捉时空相关性的强大潜力。然而,城市环境持续演变,现有模型评估多局限于交通场景,且仅使用训练后数周内数据进行测试,对模型泛化能力的探索严重不足。为此,我们提出了一个时空分布外(ST-OOD)评估基准,包含六个城市场景:共享单车、311服务、行人计数、交通速度、交通流量、网约车需求,每个场景均设置同年内(分布内)与跨年(分布外)两种情形。我们系统评估了当前主流时空模型,发现其在分布外设置下性能显著下降,多数模型表现甚至不如简单的多层感知机(MLP)。结果表明,当前领先方法过度依赖参数过拟合训练数据,导致分布内表现优异但泛化能力差。我们进一步研究了丢弃法(dropout)是否能缓解过拟合问题,发现微小的丢弃率可显著提升多数数据集上的泛化性能,对分布内性能影响极小。然而,平衡分布内与分布外表现仍是难题。我们希望该基准能推动这一关键问题的研究。

原文摘要 · Abstract (English)

Spatiotemporal neural networks have shown great promise in urban scenarios by effectively capturing temporal and spatial correlations. However, urban environments are constantly evolving, and current model evaluations are often limited to traffic scenarios and use data mainly collected only a few weeks after training period to evaluate model performance. The generalization ability of these models remains largely unexplored. To address this, we propose a Spatiotemporal Out-of-Distribution (ST-OOD) benchmark, which comprises six urban scenario: bike-sharing, 311 services, pedestrian counts, traffic speed, traffic flow, ride-hailing demand, and bike-sharing, each with in-distribution (same year) and out-of-distribution (next years) settings. We extensively evaluate state-of-the-art spatiotemporal models and find that their performance degrades significantly in out-of-distribution settings, with most models performing even worse than a simple Multi-Layer Perceptron (MLP). Our findings suggest that current leading methods tend to over-rely on parameters to overfit training data, which may lead to good performance on in-distribution data but often results in poor generalization. We also investigated whether dropout could mitigate the negative effects of overfitting. Our results showed that a slight dropout rate could significantly improve generalization performance on most datasets, with minimal impact on in-distribution performance. However, balancing in-distribution and out-of-distribution performance remains a challenging problem. We hope that the proposed benchmark will encourage further research on this critical issue.

时空模型泛化能力城市计算分布外检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。