融合政府评级与众包报告,用图神经网络更准预测城市隐患
Urban Incident Prediction with Graph Neural Networks: Integrating Government Ratings and Crowdsourced Reports
- 构建多视角多输出GNN模型,同时利用有偏众包数据和无偏政府评级
- 在纽约市数据上验证,当政府数据稀疏时仍能提升预测准确率
- 揭示高收入区报告率更高,适合城市治理与数据质量研究者
图神经网络广泛应用于城市时空预测,如基础设施问题。政府通过检查评分获取各街区真实事件状态(如道路状况),但评分覆盖稀疏、类型有限;而众包报告虽密集,却受报告行为差异影响存在偏差。本文提出一种多视角、多输出的GNN模型,融合无偏评分与有偏报告数据,以推断事件的真实潜在状态。以纽约市为例,收集并公开了3年内涵盖139类事件的961万5863条众包报告与104万1415条政府评分。实验表明,在真实及半合成数据上,本模型优于仅使用报告或仅使用评分的模型,尤其在政府数据稀疏且报告可预测评分时表现更优。此外,量化发现众包报告存在人口统计学偏差——高收入社区报告率更高。该方法为处理异构、稀疏且有偏数据下的潜在状态预测提供了普适方案。
原文摘要 · Abstract (English)
Graph neural networks (GNNs) are widely used in urban spatiotemporal forecasting, such as predicting infrastructure problems. In this setting, government officials wish to know in which neighborhoods incidents like potholes or rodent issues occur. The true state of incidents (e.g., street conditions) for each neighborhood is observed via government inspection ratings. However, these ratings are only conducted for a sparse set of neighborhoods and incident types. We also observe the state of incidents via crowdsourced reports, which are more densely observed but may be biased due to heterogeneous reporting behavior. First, for such settings, we propose a multiview, multioutput GNN-based model that uses both unbiased rating data and biased reporting data to predict the true latent state of incidents. Second, we investigate a case study of New York City urban incidents and collect, standardize, and make publicly available a dataset of 9,615,863 crowdsourced reports and 1,041,415 government inspection ratings over 3 years and across 139 types of incidents. Finally, we show on both real and semi-synthetic data that our model can better predict the latent state compared to models that use only reporting data or models that use only rating data, especially when rating data is sparse and reports are predictive of ratings. We also quantify demographic biases in crowdsourced reporting, e.g., higher-income neighborhoods report problems at higher rates. Our analysis showcases a widely applicable approach for latent state prediction using heterogeneous, sparse, and biased data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。