用多视角对比学习生成可解释的地理嵌入,提升保险风险建模精度。
A multi-view contrastive learning framework for spatial embeddings in risk modelling
- 通过卫星图与地图数据多视角对比,学习地理空间嵌入
- 在法国房价预测中提升各类模型准确率,且支持跨区域泛化
- 适合需要地理特征的保险风控、城市规划等场景
将气候、天气及人口等空间信息融入保险精算,对提高承保精度和风险管理至关重要。然而空间数据常为非结构化、高维,难以直接用于预测模型。需通过嵌入方法将其转化为有意义的表示。本文提出一种新型多视角对比学习框架,融合多源空间数据生成地理嵌入。基于欧洲范围内的卫星影像与OpenStreetMap特征构建数据集,通过坐标编码对齐不同视图,生成低维嵌入以捕捉空间结构与上下文相似性。训练后,仅需经纬度即可生成嵌入,无需原始空间输入。在法国房地产价格案例中,使用该嵌入的模型在广义线性、可加及提升模型中均显著提升预测性能,并提供可解释的空间效应;同时展现出对无训练观测区域的泛化能力。比利时洪水索赔数案例进一步验证其在保险地理风险分类中的有效性。
原文摘要 · Abstract (English)
Incorporating spatial information, particularly when related to climate, weather, and demographic factors, is crucial for improving underwriting precision and enhancing risk management in insurance. However, spatial data are often unstructured, high-dimensional, and difficult to integrate into predictive models. Embedding methods are needed to convert spatial data into meaningful representations for modelling tasks. We propose a novel multi-view contrastive learning framework for generating spatial embeddings that combine information from multiple spatial data sources. To train the model, we construct a spatial dataset that merges satellite imagery and OpenStreetMap features across Europe. The framework aligns these spatial views with coordinate-based encodings, producing low-dimensional embeddings that capture both spatial structure and contextual similarity. Once trained, the model generates embeddings directly from latitude-longitude pairs, enabling any dataset with coordinates to be enriched with meaningful spatial features without requiring access to the original spatial inputs. In a case study on French real estate prices, we compare models trained on raw coordinates against those using our spatial embeddings as inputs. The embeddings consistently improve predictive accuracy across generalised linear, additive, and boosting models, while providing post-hoc explainable spatial effects and demonstrating generalisation of the fitted spatial effects to regions without training observations. A second case study on flood claim counts across Belgian postal codes confirms that the embeddings improve territorial risk classification in an insurance context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。