arXiv:2606.23453cs.LG2026-06

测试11种位置编码在不同尺度下能否准确还原空间效应。

Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales

论文配图:Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales
图 1 · 摘自论文原文
  • 用博弈论方法评估位置编码对空间效应的捕捉能力
  • 全球尺度下次要系数恢复效果最差,主系数始终表现良好
  • 原始坐标在所有场景中仍具竞争力,适合作为基线

位置编码将地理坐标转换为高维嵌入以支持下游机器学习任务,但其对可解释空间效应的捕捉能力尚不明确。本文采用基于博弈论的GeoShapley解释器,将所有位置特征视为单一联合参与者,评估其能否从位置编码嵌入构建的模型中恢复出已知的空间变系数。在三个尺度(网格、县、全球)下,对比了11种来自TorchSpatial框架的位置编码,在有无原始坐标的条件下,以及未训练与对比训练两种情形下的表现。通过估计系数与真实系数的相关性衡量恢复效果。结果显示,主系数的恢复率在各编码器中保持高位,而次级系数的恢复受尺度影响显著,尤其在全局尺度差异最大;原始坐标基线在所有条件下均表现稳健,具有较强竞争力。

原文摘要 · Abstract (English)

Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these representations capture interpretable spatial effects. We benchmark whether GeoShapley, a game-theoretic explainer that treats all location features as a single joint player, can recover spatially varying coefficients from models built on location-encoder embeddings. Eleven encoders from the TorchSpatial framework are evaluated against a synthetic process with known coefficients, across three scales (grid, county, global), with and without raw coordinates alongside the embedding, and under untrained and contrastively trained conditions. Measuring recovery as the correlation between estimated and true coefficients, we report how it varies with scale and encoder architecture and compare the embeddings against a raw-coordinate baseline. Recovery of the primary coefficient is consistently high across encoders, whereas recovery of a secondary coefficient is more scale-dependent, differing most at the global scale; the raw-coordinate baseline remains competitive throughout.

位置编码空间建模可解释性地理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。