用表征学习增强地震数据插值,提升缺失数据重建精度与效率
SeisRDT: Latent Diffusion Model Based On Representation Learning For Seismic Data Interpolation And Reconstruction
- 基于表征学习的掩码建模,利用已知数据推断缺失数据
- 在真实与合成数据上均优于现有方法,支持复杂缺失模式
- 结合预训练压缩模型,显著降低计算开销,适合工业级应用
由于地理、物理或经济因素限制,采集的地震数据常存在缺失道。传统重建方法需手动调参,难以处理大规模连续缺失。深度学习中各类扩散模型虽具强重建能力,但基于UNet的架构计算量大,且难捕捉地震数据各道间的相关性。为应对地震数据中复杂不规则的缺失情况,本文提出一种基于表征学习的潜在扩散变压器(SeisRDT)。通过表征学习的掩码建模机制,利用已知数据的标记序列推断未知数据的标记序列,使扩散模型生成的数据分布更一致,相关性与已知数据更匹配。设计了带有相对位置偏置的注意力机制,实现对地震数据的全局建模。采用预训练数据压缩模型将扩散模型的训练与推理过程映射至潜空间,相比其他基于扩散模型的方法显著降低计算与推理成本。在野外及合成数据集上的重建实验表明,本方法在多种复杂缺失场景下均取得更高重建精度。
原文摘要 · Abstract (English)
Due to limitations such as geographic, physical, or economic factors, collected seismic data often have missing traces. Traditional seismic data reconstruction methods face the challenge of selecting numerous empirical parameters and struggle to handle large-scale continuous missing traces. With the advancement of deep learning, various diffusion models have demonstrated strong reconstruction capabilities. However, these UNet-based diffusion models require significant computational resources and struggle to learn the correlation between different traces in seismic data. To address the complex and irregular missing situations in seismic data, we propose a latent diffusion transformer utilizing representation learning for seismic data reconstruction. By employing a mask modeling scheme based on representation learning, the representation module uses the token sequence of known data to infer the token sequence of unknown data, enabling the reconstructed data from the diffusion model to have a more consistent data distribution and better correlation and accuracy with the known data. We propose the Representation Diffusion Transformer architecture, and a relative positional bias is added when calculating attention, enabling the diffusion model to achieve global modeling capability for seismic data. Using a pre-trained data compression model compresses the training and inference processes of the diffusion model into a latent space, which, compared to other diffusion model-based reconstruction methods, reduces computational and inference costs. Reconstruction experiments on field and synthetic datasets indicate that our method achieves higher reconstruction accuracy than existing methods and can handle various complex missing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。