首个可跨地震调查通用解释任意地质体的多模态基础模型
A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys
- 用预训练视觉模型+多模态提示引擎实现跨调查泛化
- 2D到3D扩展性强,未见地质体也能准确识别
- 支持实时交互,适合地质勘探与工程应用
地震地质体解释对构造地质研究和各类工程应用至关重要。现有深度学习方法虽有潜力,但缺乏多模态输入支持,且难以泛化到不同地质体或调查区域。本文提出一种可提示的基础模型,用于跨地震调查解释任意地质体。该模型融合预训练视觉基础模型(VFM)与复杂多模态提示引擎。VFM在大量自然图像上预训练,并在地震数据上微调,提供跨调查泛化的鲁棒特征提取能力。提示引擎结合多模态先验信息,迭代优化地质体边界。大量实验表明,该模型在准确性、从2D到3D的可扩展性以及对多种地质体(包括训练中未见类型)的泛化能力方面均表现优异。据我们所知,这是首个高度可扩展且多功能的多模态基础模型,能跨调查解释任意地质体并支持实时交互。本方法为地球科学数据解释树立了新范式,具有向其他任务迁移的广泛潜力。
原文摘要 · Abstract (English)
Seismic geobody interpretation is crucial for structural geology studies and various engineering applications. Existing deep learning methods show promise but lack support for multi-modal inputs and struggle to generalize to different geobody types or surveys. We introduce a promptable foundation model for interpreting any geobodies across seismic surveys. This model integrates a pre-trained vision foundation model (VFM) with a sophisticated multi-modal prompt engine. The VFM, pre-trained on massive natural images and fine-tuned on seismic data, provides robust feature extraction for cross-survey generalization. The prompt engine incorporates multi-modal prior information to iteratively refine geobody delineation. Extensive experiments demonstrate the model's superior accuracy, scalability from 2D to 3D, and generalizability to various geobody types, including those unseen during training. To our knowledge, this is the first highly scalable and versatile multi-modal foundation model capable of interpreting any geobodies across surveys while supporting real-time interactions. Our approach establishes a new paradigm for geoscientific data interpretation, with broad potential for transfer to other tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。