用4D时空编码构建全球尺度自监督世界模型,实现厘米级秒级精准预测。
Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding
- 引入地球级4D时空位置编码,融合时间维度扩展3D哈希编码。
- 在生态预测任务上超越更大规模预训练多模态模型,达顶尖性能。
- 适合遥感、气候建模与长期预测研究者,开源可复现。
我们提出DeepEarth,一种基于Earth4D的自监督多模态世界模型。Earth4D是一种新型行星尺度4D时空位置编码器,将3D多分辨率哈希编码扩展至包含时间维度,在跨世纪尺度上实现亚米级、亚秒级精度的高效扩展。多模态编码器(如视觉-语言模型)与Earth4D嵌入融合,并通过掩码重建进行训练。我们在生态预测基准上验证了Earth4D的强大表达能力,其表现优于在更大数据集上预训练的多模态基础模型。代码与模型已开源:https://github.com/legel/deepearth
原文摘要 · Abstract (English)
We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。