arXiv:2602.04890physics.geo-phcs.AI2026-02被引 1

构建首个覆盖广泛地质条件的2D地震图像数据集,助力机器学习模型训练与泛化评估。

A General-Purpose Diversified 2D Seismic Image Dataset from NAMSS

  • 从全球122个海域采集2588张标准化地震剖面,覆盖多区域、多采样条件。
  • 采用区域不重叠划分,确保模型在未见地质条件下仍具泛化能力。
  • 相比现有数据集,覆盖更广的地震图像空间,适合预训练与迁移学习。

我们推出了Unicamp-NAMSS数据集,这是一个大规模、多样化且地理分布广泛的迁移后2D地震剖面集合,旨在支持地球物理学中的现代机器学习研究。该数据集源自国家海洋地震调查档案(NAMSS),包含数十年来公开的全球海洋地震数据,涵盖多个区域、采集条件和地质环境。经过全面收集与筛选,共获得2588张清洗并标准化的地震剖面,来自122个调查区域,覆盖广泛的垂直与水平采样特性。为确保实验可靠性,我们对数据集进行了平衡处理,避免单一调查主导分布,并将其划分为互不重叠的宏观区域用于训练、验证和测试。这种区域分离的划分方式可实现对未见地质与采集条件的稳健泛化评估。通过卷积与Transformer模型的定量分析与嵌入空间检验,验证了Unicamp-NAMSS在区域内及跨区域均具有显著变异性,同时保持了采集宏观区域与调查类型间的结构一致性。与广泛应用的解释数据集(Parihaka与F3 Block)对比显示,Unicamp-NAMSS覆盖了更广阔的地震外观空间,是机器学习模型预训练的理想候选。因此,该数据集为自监督表示学习、迁移学习、超分辨率或属性预测等监督任务的基准测试,以及地震解释中的域适应研究提供了宝贵资源。

原文摘要 · Abstract (English)

We introduce the Unicamp-NAMSS dataset, a large, diverse, and geographically distributed collection of migrated 2D seismic sections designed to support modern machine learning research in geophysics. We constructed the dataset from the National Archive of Marine Seismic Surveys (NAMSS), which contains decades of publicly available marine seismic data acquired across multiple regions, acquisition conditions, and geological settings. After a comprehensive collection and filtering process, we obtained 2588 cleaned and standardized seismic sections from 122 survey areas, covering a wide range of vertical and horizontal sampling characteristics. To ensure reliable experimentation, we balanced the dataset so that no survey dominates the distribution, and partitioned it into non-overlapping macro-regions for training, validation, and testing. This region-disjoint split allows robust evaluation of generalization to unseen geological and acquisition conditions. We validated the dataset through quantitative and embedding-space analyses using both convolutional and transformer-based models. These analyses showed that Unicamp-NAMSS exhibits substantial variability within and across regions, while maintaining coherent structure across acquisition macro-region and survey types. Comparisons with widely used interpretation datasets (Parihaka and F3 Block) further demonstrated that Unicamp-NAMSS covers a broader portion of the seismic appearance space, making it a strong candidate for machine learning model pretraining. The dataset, therefore, provides a valuable resource for machine learning tasks, including self-supervised representation learning, transfer learning, benchmarking supervised tasks such as super-resolution or attribute prediction, and studying domain adaptation in seismic interpretation.

地震图像数据集机器学习泛化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。