开源真实地震与井数据集,助力生成式AI做地震反演
OpenSeisML: Open Large-Scale Real Seismic and well-log Dataset for Generative AI

- 从公开数据构建地震与井数据联动的时深转换模型
- 提供可复现的数据处理流程,支持生成模型训练
- 适合做地质建模与不确定性量化研究的团队使用
机器学习与计算机视觉的发展显著降低了传统迭代式地震反演的计算成本。然而,高质量速度模型的稀缺限制了机器学习方法的开发与评估,因多数优质数据由油气公司私有。为此,我们推出OpenSeisML,一个面向生成式AI(Gen-AI)地震反演的大型真实地震与井数据集。数据源自英国国家数据存档库(NDR)公开调查资料。当地震数据为时间域、井数据为深度域时,需进行时深转换。我们利用井中检波器数据建立时深关系,并通过插值构建速度模型,实现叠后地震数据的准确转换。本文提出自动化数据整理流程,保障可复现性。目标是训练生成模型,捕捉地下属性的统计分布,合成多个统计一致的实现,用于不确定性量化,并作为地震反演的先验信息。
原文摘要 · Abstract (English)
The advent of machine learning (ML) and computer vision has significantly accelerated seismic inversion workflows by reducing the computational cost of traditionally expensive iterative methods. However, the development and evaluation of ML methods remain limited by the scarcity of realistic velocity models, as most high-quality data are privately owned by oil and gas companies. To address this gap, we present OpenSeisML, a collection of real seismic datasets designed to support generative AI (Gen-AI) workflows for seismic inversion. The datasets are curated from publicly available surveys in the UK National Data Repository (NDR). When seismic volumes are in the time domain and wells are in depth, a time-to-depth conversion is required. We use checkshot data to establish the time-depth relationship and construct a velocity model through interpolation for accurate conversion of post-stack seismic data. Here, we present an automated data curation pipeline that enables seismic data preparation while ensuring reproducibility. The objective is to train a generative model that captures the statistical distribution of subsurface properties, enabling the synthesis of multiple statistically consistent realizations for uncertainty quantification which can act as a prior for seismic inversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。