用自监督学习提升地质XRF数据建模精度,仅需1/3数据就超越传统方法。
MAX: Masked Autoencoder for X-ray Fluorescence in Geological Investigation
- 设计掩码自编码器MAX,对高分辨率XRF光谱进行自监督预训练。
- 在钙碳酸盐和有机碳定量任务中,仅用1/3数据即达到更高准确率。
- 零样本测试下泛化能力提升60%以上,适合缺乏标注数据的地质研究者。
预训练基础模型已成为深度学习的标准流程,但在地质学中应用仍受限,主要因数据稀缺导致模型迁移能力不足。本文聚焦于科学钻探项目中的标准高分辨率测量——X射线荧光(XRF)光谱数据,提出可扩展的自监督学习框架MAX,用于构建覆盖太平洋与南大洋多个区域地质记录的基础模型。预训练阶段发现,对输入光谱进行50%的掩码可生成有意义的自监督任务。下游任务选择两种成本高昂的地球化学指标:碳酸钙(CaCO₃)和总有机碳(TOC),以理解古海洋碳循环。结果表明,仅需1/3的数据,MAX在量化任务上的准确率已优于无预训练模型;且在新样品的零样本测试中,模型泛化能力提升超过60%,可解释性进一步增强了其鲁棒性。该方法为克服地质发现中的数据稀缺问题提供了可行路径。
原文摘要 · Abstract (English)
Pre-training foundation models has become the de-facto procedure for deep learning approaches, yet its application remains limited in the geological studies, where in needs of the model transferability to break the shackle of data scarcity. Here we target on the X-ray fluorescence (XRF) scanning data, a standard high-resolution measurement in extensive scientific drilling projects. We propose a scalable self-supervised learner, masked autoencoders on XRF spectra (MAX), to pre-train a foundation model covering geological records from multiple regions of the Pacific and Southern Ocean. In pre-training, we find that masking a high proportion of the input spectrum (50\%) yields a nontrivial and meaningful self-supervisory task. For downstream tasks, we select the quantification of XRF spectra into two costly geochemical measurements, CaCO$_3$ and total organic carbon, due to their importance in understanding the paleo-oceanic carbon system. Our results show that MAX, requiring only one-third of the data, outperforms models without pre-training in terms of quantification accuracy. Additionally, the model's generalizability improves by more than 60\% in zero-shot tests on new materials, with explainability further ensuring its robustness. Thus, our approach offers a promising pathway to overcome data scarcity in geological discovery by leveraging the self-supervised foundation model and fast-acquired XRF scanning data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。