arXiv:2607.29527cs.LGcs.AI2026-07

构建跨地球与社会的统一模型,实现多源异构数据融合与高效推理。

TerraNova: A Foundation Model for the Anthropocene

论文配图:TerraNova: A Foundation Model for the Anthropocene
图 1 · 摘自论文原文
  • 采用原生几何编码,分别处理网格化地球场与国家指标数据
  • 在1024个数据集上训练,支持时间、海洋与不确定性建模
  • 可快速适配新变量,单卡消费级硬件分钟级完成微调

人类世的核心挑战在于将物理地球与人类社会作为耦合系统建模,但现有学习表征无法覆盖其观测广度。我们指出障碍源于几何差异:地球系统以连续场形式测量,忽略政治边界;而社会数据按行政单元报告。传统方法需对边界进行有损平均。本文提出TerraNova,一个在1,024个原始几何数据上训练的基础模型,包含512个网格化地球系统场与512个国家指标。专用编码器分别表示位置、国家、时间与任务,跨模态变换器将其融合为共享时空状态,超网络生成每查询专属解码器,其证据头输出预测分布。通过两种对比目标耦合表征:基于人口加权的国家-坐标对齐,以及与预训练地理空间嵌入的语义一致性。该模型读出后性能媲美专用地理编码器,同时覆盖其未涵盖维度(时间、海洋、不确定性),并支持国家层面能力。冻结主干可从稀疏观测重建密集场,且在消费级硬件上几分钟内即可适应未见变量。

原文摘要 · Abstract (English)

A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representation: a population-weighted alignment between each country and coordinates in its territory, and one to pretrained geospatial embeddings carrying image-derived semantics. Read out through that decoder, the representation is competitive with purpose-built geospatial encoders while spanning axes they do not represent (time, oceans and uncertainty) and supporting country-level capabilities. The frozen backbone reconstructs dense fields from sparse observations and adapts to unseen variables in minutes on consumer hardware.

基础模型地球系统多模态时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。