用双网格与希尔伯特编码提升地震图像视觉大模型预训练效果
Synergizing Multigrid Algorithms with Vision Transformer: A Novel Approach to Enhance the Seismic Foundation Model
- 分频处理+分层希尔伯特编码,兼顾地震数据高频低频特征
- 自适应训练策略先学粗结构后调细特征,提升模型性能
- 专为地震波形设计,适合地震领域视觉大模型研发者
由于人工智能技术的快速发展,基于Transformer的视觉基础模型已在药物发现、材料研究和天文学等领域取得突破。然而,地震数据具有独特性,其高、低频成分对预训练至关重要。现有视觉变压器(ViT)采用顺序标记化,忽略内在结构,难以有效捕捉地震数据中的高低频信息。本文提出一种针对地震波形数据的自适应双网格基础模型训练策略(ADATG),结合谱分解分离高低频分量,并利用分层希尔伯特编码进行高效表示。此外,基于ViT中观察到的频率规律,设计了一种自适应训练机制:初期聚焦粗粒度信息,随后逐步增强对细粒度特征的学习。大量实验验证了该方法的有效性和效率。研究表明,结合地震数据特性设计的数据编码与训练策略对视觉地震基础模型的预训练具有关键意义。
原文摘要 · Abstract (English)
Due to the emergency and homogenization of Artificial Intelligence (AI) technology development, transformer-based foundation models have revolutionized scientific applications, such as drug discovery, materials research, and astronomy. However, seismic data presents unique characteristics that require specialized processing techniques for pretraining foundation models in seismic contexts with high- and low-frequency features playing crucial roles. Existing vision transformers (ViTs) with sequential tokenization ignore the intrinsic pattern and fail to grasp both the high- and low-frequency seismic information efficiently and effectively. This work introduces a novel adaptive two-grid foundation model training strategy (ADATG) with Hilbert encoding specifically tailored for seismogram data, leveraging the hierarchical structures inherent in seismic data. Specifically, our approach employs spectrum decomposition to separate high- and low-frequency components and utilizes hierarchical Hilbert encoding to represent the data effectively. Moreover, observing the frequency principle observed in ViTs, we propose an adaptive training strategy that initially emphasizes coarse-level information and then progressively refines the model's focus on fine-level features. Our extensive experiments demonstrate the effectiveness and efficiency of our training methods. This research highlights the importance of data encoding and training strategies informed by the distinct characteristics of high- and low-frequency features in seismic images, ultimately contributing to the enhancement of visual seismic foundation models pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。