用变分自编码器高效建模区域间空间依赖,提升小区域估计速度与可扩展性。
Variational Autoencoded Multivariate Spatial Fay-Herriot Models
- 用变分自编码器学习空间随机效应,捕捉区域间依赖结构
- 训练一次后可复用,大幅降低后续建模计算开销
- 在加州全普查区数据上验证,处理大规模多变量数据高效可靠
小区域估计模型对样本量有限区域的人口特征推断至关重要,支持政策制定、人口研究和资源分配。空间费希尔-赫里奥特模型通过引入空间相关性,利用邻近区域信息提升估计精度,但计算开销大,难以扩展至高维多变量数据。本文提出两种方法,将多变量空间费希尔-赫里奥特模型与通过变分自编码器学习的空间随机效应结合,高效利用空间结构。重要的是,训练完成后,该编码器可在后续建模中重复使用,无需重新训练。此外,利用变分自编码器表示空间依赖显著提升计算效率,适用于大规模数据集。我们在美国社区调查5年期数据(覆盖加州所有普查区)上验证了方法的有效性。
原文摘要 · Abstract (English)
Small area estimation models are essential for estimating population characteristics in regions with limited sample sizes, thereby supporting policy decisions, demographic studies, and resource allocation, among other use cases. The spatial Fay-Herriot model is one such approach that incorporates spatial dependence to improve estimation by borrowing strength from neighboring regions. However, this approach often requires substantial computational resources, limiting its scalability for high-dimensional datasets, especially when considering multiple (multivariate) responses. This paper proposes two methods that integrate the multivariate spatial Fay-Herriot model with spatial random effects, learned through variational autoencoders, to efficiently leverage spatial structure. Importantly, after training the variational autoencoder to represent spatial dependence for a given set of geographies, it may be used again in future modeling efforts, without the need for retraining. Additionally, the use of the variational autoencoder to represent spatial dependence results in extreme improvements in computational efficiency, even for massive datasets. We demonstrate the effectiveness of our approach using 5-year period estimates from the American Community Survey over all census tracts in California.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。