首个面向月球遥感的多模态基础模型与评测基准,打通多源数据孤岛。
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing

- 构建28通道月球预训练数据集,覆盖7类仪器、5个任务,分辨率约237米/像素
- 提出MG-MAE模型,支持缺失模态、异质空间覆盖和物理可解释重建,性能超越基线
- 提供6项下游任务评测,适用于月球地质、资源探测等研究者
数十年轨道任务积累了涵盖光学影像、光谱、热辐射、雷达、重力和元素组成的多模态月球遥感数据,但这些数据分散在不同档案中,且缺乏机器学习评估基准。我们提出Moonstone,首个面向月球遥感的多模态基础模型与评测基准。贡献包括:(1)来自五个任务中七类仪器的28通道、128像素/度(约237米/像素)全球预训练数据集;(2)一种基于模态分组的掩码自编码器MG-MAE,包含每组独立卷积令牌化器、共享视觉变压器编码器、缺失模态注意力掩码、覆盖自适应掩码及光谱连续性正则化,以实现物理合理的重建;(3)涵盖分类、回归与分割的六项下游任务基准。在所有任务上,预训练特征均优于随机初始化基线,并显著超越ImageNet预训练与原始MAE基线。数据与代码已公开于https://huggingface.co/datasets/ayushprd/Moonstone 和 https://github.com/ayushprd/Moonstone。
原文摘要 · Abstract (English)
Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emission, radar, gravity, and elemental composition. Yet these datasets remain fragmented across archives, and no benchmark exists for evaluating machine learning on lunar data. We introduce Moonstone, the first multi-modal foundation model benchmark for lunar remote sensing. Our contributions are: (1) a 28-channel, 128 pixels-per-degree (~237 m) global lunar pretraining dataset from seven instrument families across five missions, (2) MG-MAE, a modality-grouped masked autoencoder with per-group convolutional tokenizers, a shared Vision Transformer encoder, attention masking for missing modalities, coverage-adaptive masking for heterogeneous spatial coverage, and spectral continuity regularization for physically plausible reconstructions, and (3) a benchmark of six downstream tasks covering classification, regression, and segmentation. MG-MAE pretrained features outperform scratch baselines on all tasks and surpass both ImageNet-pretrained and vanilla MAE baselines by large margins. Data and code are available at https://huggingface.co/datasets/ayushprd/Moonstone and https://github.com/ayushprd/Moonstone .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。