用搜索频率生成社区级特征,提升地理空间预测精度。
Community search signatures as foundation features for human-centered geospatial modeling
- 基于社区搜索频次构建匿名化基础特征。
- 在95%美国人口覆盖区,健康变量预测平均R²达0.74。
- 无需严格时间对齐,优于传统插值与卫星图像模型。
聚合的相对搜索频率提供了一种独特复合信号,反映人们的行为习惯、关注点、兴趣、意图及信息需求,这是其他现成数据集无法提供的。时间搜索趋势已在传染病、失业率、零售销售等多个领域成功用于时间序列建模。然而,现有方法大多依赖专门整理的关键词、查询或查询聚类数据集,且搜索数据需与目标变量进行时间对齐。本文提出一种新方法,生成社区层级的聚合与匿名化搜索兴趣表示,作为地理空间建模的基础特征。我们在多个领域的空间数据集上进行了基准测试。在人口超过3000的邮编区(覆盖美国本土95%以上人口)中,针对20%保留县份的缺失值预测,模型在21个健康变量上的平均R²达到0.74,在6个人口与环境变量上达到0.80。结果表明,这些搜索特征可用于空间预测而无需严格时间对齐,且性能优于空间插值和使用卫星图像特征的前沿方法。
原文摘要 · Abstract (English)
Aggregated relative search frequencies offer a unique composite signal reflecting people's habits, concerns, interests, intents, and general information needs, which are not found in other readily available datasets. Temporal search trends have been successfully used in time series modeling across a variety of domains such as infectious diseases, unemployment rates, and retail sales. However, most existing applications require curating specialized datasets of individual keywords, queries, or query clusters, and the search data need to be temporally aligned with the outcome variable of interest. We propose a novel approach for generating an aggregated and anonymized representation of search interest as foundation features at the community level for geospatial modeling. We benchmark these features using spatial datasets across multiple domains. In zip codes with a population greater than 3000 that cover over 95% of the contiguous US population, our models for predicting missing values in a 20% set of holdout counties achieve an average $R^2$ score of 0.74 across 21 health variables, and 0.80 across 6 demographic and environmental variables. Our results demonstrate that these search features can be used for spatial predictions without strict temporal alignment, and that the resulting models outperform spatial interpolation and state of the art methods using satellite imagery features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。