arXiv:2507.13323cs.LG2025-07被引 1

用大模型辅助少样本回归,精准估算欠发达地区经济社会指标

GeoReg: Weight-Constrained Few-Shot Regression for Socio-Economic Estimation using LLM

  • 借大模型提取特征,构建带约束的线性估计器
  • 在低收入国家数据稀缺下仍实现高精度预测
  • 适合政策研究与资源分配场景使用

区域生产总值、人口、教育水平等社会经济指标对政策制定和可持续发展至关重要。本文提出GeoReg,一种融合卫星图像与网络地理空间信息的回归模型,可在数据匮乏地区(如发展中国家)进行有效估计。该方法利用大语言模型的先验知识,使其充当数据工程师角色,提取有信息量的特征以支持少样本学习。具体而言,模型分析特征与目标指标之间的上下文关系,将其分类为正相关、负相关、混合或无关,并针对每类设置特定权重约束。为捕捉非线性模式,模型还识别有意义的特征交互并整合非线性变换。在三个不同发展阶段国家的实验表明,该模型在少样本条件下显著优于基线方法,尤其在低收入国家数据有限时表现优异。

原文摘要 · Abstract (English)

Socio-economic indicators like regional GDP, population, and education levels, are crucial to shaping policy decisions and fostering sustainable development. This research introduces GeoReg a regression model that integrates diverse data sources, including satellite imagery and web-based geospatial information, to estimate these indicators even for data-scarce regions such as developing countries. Our approach leverages the prior knowledge of large language model to address the scarcity of labeled data, with the language model functioning as a data engineer by extracting informative features to enable effective estimation in few-shot settings. Specifically, our model obtains contextual relationships between data features and the target indicator, categorizing their correlations as positive, negative, mixed, or irrelevant. These features are then fed into the linear estimator with tailored weight constraints for each category. To capture nonlinear patterns, the model also identifies meaningful feature interactions and integrates them, along with nonlinear transformations. Experiments across three countries at different stages of development demonstrate that our model outperforms baselines in estimating socio-economic indicators, even for low-income countries with limited data availability.

少样本学习社会经济估算大模型应用地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。