arXiv:2507.19697cs.LG2025-07

用行业分类信息提升大规模门店共访预测准确率

NAICS-Aware Graph Neural Networks for Large-Scale POI Co-visitation Prediction: A Multi-Modal Dataset and Methodology

  • 引入行业代码嵌入,让图神经网络理解门店间语义关系
  • 在42亿潜在组合中实现0.625的R²,提升157%
  • 适合城市规划、零售分析等需要精准客流预测的场景

理解人们访问一家商家后会去哪些地方,对城市规划、零售分析和位置服务至关重要。然而,由于数据极度稀疏且空间距离与商业关系交织,跨百万级商家预测共访模式仍具挑战。传统仅依赖地理距离的方法无法解释为何咖啡馆与高端餐厅即便毗邻也吸引不同客流。我们提出NAICS-aware GraphSAGE,一种融合行业分类知识的可学习嵌入的图神经网络,用于预测大规模共访模式。核心洞察是:通过详细行业代码捕捉的商业语义,能提供纯空间模型无法解释的关键信号。该方法通过按州分解高效处理超大规模数据(42亿潜在门店对),并端到端整合空间、时间与社会经济特征。在包含9490万条共访记录、92,486个品牌及48个美国州的POI-Graph数据集上评估,相比最先进基线,R²从0.243提升至0.625(提升157%),排名质量(NDCG@10)提升32%。

原文摘要 · Abstract (English)

Understanding where people go after visiting one business is crucial for urban planning, retail analytics, and location-based services. However, predicting these co-visitation patterns across millions of venues remains challenging due to extreme data sparsity and the complex interplay between spatial proximity and business relationships. Traditional approaches using only geographic distance fail to capture why coffee shops attract different customer flows than fine dining restaurants, even when co-located. We introduce NAICS-aware GraphSAGE, a novel graph neural network that integrates business taxonomy knowledge through learnable embeddings to predict population-scale co-visitation patterns. Our key insight is that business semantics, captured through detailed industry codes, provide crucial signals that pure spatial models cannot explain. The approach scales to massive datasets (4.2 billion potential venue pairs) through efficient state-wise decomposition while combining spatial, temporal, and socioeconomic features in an end-to-end framework. Evaluated on our POI-Graph dataset comprising 94.9 million co-visitation records across 92,486 brands and 48 US states, our method achieves significant improvements over state-of-the-art baselines: the R-squared value increases from 0.243 to 0.625 (a 157 percent improvement), with strong gains in ranking quality (32 percent improvement in NDCG at 10).

图神经网络共访预测多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。