arXiv:2501.00941cs.LGcs.CV2025-01AAAI被引 6

针对地质数据不平衡问题,提出新型扩散模型生成成对科学数据。

A Novel Diffusion Model for Pairwise Geoscience Data Generation with Unbalanced Training Dataset

  • 设计一进二出编码器-解码器结构,从共享潜在空间生成成对数据。
  • 在OpenFWI数据集上FID得分显著优于现有方法,生成数据质量更高。
  • 适合处理地震成像等多模态数据稀缺且不均衡的科学计算场景。

生成式AI虽已深刻影响日常生活,但在科学领域应用仍处于起步阶段。数据稀缺是数据驱动科学计算的主要障碍,因此物理引导的生成式AI具有重要潜力。科学计算中多数任务需转换多种数据模态以描述物理现象,如地震成像中的空间与波形、信号处理中的时域与频域、气候建模中的时序与谱信息;因此,相较于自然图像中常见的单模态生成,多模态成对数据生成更为迫切。然而现实中,各模态数据常存在数量不平衡问题,例如地震成像中速度图可轻松模拟,但真实地震波形数据却严重缺乏。尽管近期研究已将强大扩散模型用于多模态生成,如何有效利用不平衡数据仍不明确。本文以地下地球物理学中的地震成像为例,提出"UB-Diff"——一种新型扩散模型,用于多模态成对科学数据生成。核心创新在于采用一进二出编码器-解码器网络结构,确保成对数据由共享潜在表示生成,随后该潜在表示被用于扩散过程生成成对数据。在OpenFWI数据集上的实验结果表明,UB-Diff在弗雷谢特起始距离(FID)评分和成对评估中均显著优于现有技术,证明其能生成可靠且有用的多模态成对数据。

原文摘要 · Abstract (English)

Recently, the advent of generative AI technologies has made transformational impacts on our daily lives, yet its application in scientific applications remains in its early stages. Data scarcity is a major, well-known barrier in data-driven scientific computing, so physics-guided generative AI holds significant promise. In scientific computing, most tasks study the conversion of multiple data modalities to describe physical phenomena, for example, spatial and waveform in seismic imaging, time and frequency in signal processing, and temporal and spectral in climate modeling; as such, multi-modal pairwise data generation is highly required instead of single-modal data generation, which is usually used in natural images (e.g., faces, scenery). Moreover, in real-world applications, the unbalance of available data in terms of modalities commonly exists; for example, the spatial data (i.e., velocity maps) in seismic imaging can be easily simulated, but real-world seismic waveform is largely lacking. While the most recent efforts enable the powerful diffusion model to generate multi-modal data, how to leverage the unbalanced available data is still unclear. In this work, we use seismic imaging in subsurface geophysics as a vehicle to present ``UB-Diff'', a novel diffusion model for multi-modal paired scientific data generation. One major innovation is a one-in-two-out encoder-decoder network structure, which can ensure pairwise data is obtained from a co-latent representation. Then, the co-latent representation will be used by the diffusion process for pairwise data generation. Experimental results on the OpenFWI dataset show that UB-Diff significantly outperforms existing techniques in terms of Fréchet Inception Distance (FID) score and pairwise evaluation, indicating the generation of reliable and useful multi-modal pairwise data.

扩散模型地质数据成对生成数据不平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。