arXiv:2508.01109cs.AI2025-08被引 4

用卫星图和网络文本预测贫困水平,发现语言与视觉信息可互补提升准确率。

Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?

  • 融合卫星图像与LLM生成/检索的文本信息进行贫困预测
  • 多模态融合使预测相关系数达0.77,显著优于仅用图像的0.63
  • 揭示语言与视觉特征存在部分对齐,但未证实单一共享表征

我们探究社会经济指标(如家庭财富)是否在卫星影像(反映建筑、道路等特征)和互联网文本(体现社区历史、文化与叙事)中留下可恢复的信息痕迹。基于非洲地区60,000个DHS集群的数据,将高分辨率Landsat图像与由条件于地理位置/年份的LLM生成的文本描述,以及由LLM驱动的AI搜索代理从网络获取的文本配对。构建五种预测管道:(i)仅使用卫星图像的视觉模型,(ii)仅用位置和年份的LLM,(iii)搜索并合成网络文本的AI代理,(iv)联合图像-文本编码器,(v)所有信号的集成。结果表明:融合视觉与代理/LLM生成文本可显著提升财富预测性能(例如,在外部样本上R²达0.77,优于仅用图像的0.63),且LLM内部知识在跨国家/时间泛化中表现意外有效;同时发现视觉与语言嵌入经对齐后中位余弦相似度约为0.60,支持柏拉图表征假说的初步证据,但未证明收敛至单一共享潜在表征;由于代理检索数据增益微弱且不稳定,对代理诱导新异性的证据有限。研究发布包含约60,000个集群的多模态数据集,每条数据关联卫星图、LLM生成描述及代理检索文本。

原文摘要 · Abstract (English)

We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery (capturing features like buildings and roads) and Internet-sourced text (reflecting historical, cultural, and narratives of neighborhoods). Using DHS data from African neighborhoods (clusters), we pair high-resolution Landsat images with textual descriptions generated by LLMs conditioned on location/year, plus text retrieved by an LLM-driven AI Search Agent from web sources. We develop a multimodal framework that predicts household wealth (International Wealth Index; IWI) via five pipelines: (i) a vision model on satellite images, (ii) an LLM using only location and year, (iii) an AI agent that searches and synthesizes web text, (iv) a joint image-text encoder, and (v) an ensemble of all signals. Our framework yields three contributions. First, evaluations show that fusing vision and agent/LLM-generated text improves on vision-only baselines in wealth prediction (e.g., R-squared of 0.77 vs. 0.63 on out-of-sample splits), with LLM-internal knowledge (artificial neural memory) proving surprisingly predictive in out-of-country/time generalization. Second, we find suggestive evidence of partial representational alignment: fused embeddings from vision and language modalities correlate moderately (median cosine similarity across modalities of about 0.60 after alignment). This pattern is broadly consistent with the Platonic Representation Hypothesis, but does not by itself establish convergence to a single shared latent representation. Because agent-retrieved data yields only marginal and unstable gains across splits, our evidence for the Agent-Induced Novelty Hypothesis is limited. Third, we release a large-scale multimodal dataset of about 60,000 DHS clusters, each linked to satellite images, LLM-generated descriptions, and AI-agent-retrieved texts.

贫困预测多模态学习大模型应用卫星图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。