构建首个地震多模态数据集与模型,实现波形、图像与文本的联合分析。
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

- 整合16000+地震事件的波形、强度图与人口暴露可视化数据
- 提出专用时序编码器,使模型在跨模态推理任务中性能提升显著
- 适合地震研究、多模态模型开发及科学领域通用模型应用者
通用多模态模型在专业科学领域的应用受限于缺乏整合多种模态的数据集。地震学需融合时间序列波形数据、地理图像和上下文元数据,而现有数据集未涵盖此多模态整合。本文提出MultiSeismo,一个大规模结构化多模态地震数据集,包含2010至2023年共13年间超过16,000个地震事件,覆盖多样地理区域。每条事件数据集成全球台站网络的波形记录、强度图、人口暴露可视化及标准化JSON格式的文本描述。我们进一步构建MISCE指令集,支持通用多模态模型在地震推理任务(从信息检索到复杂跨模态分析)上的监督训练与评估。基于MISCE微调统一输入模型(Unified IO 2)并引入专用时序编码器,得到首个针对地震学的多模态模型SeisModal。对现有先进多模态模型在MultiSeismo上的评估显示,通用模型在处理时间序列数据上存在显著挑战,而SeisModal在地震多模态推理任务中表现更优。结果表明,MultiSeismo为地震学多模态研究提供严格基准,并验证了领域特化架构的有效性。
原文摘要 · Abstract (English)
The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-specific datasets that integrate multiple data modalities beyond text and images. In seismology, understanding earthquake phenomena requires the synthesis of timeseries waveform data, geographical imagery, and contextual metadata, a multimodal integration absent in existing seismic datasets. We present MultiSeismo, a large scale structured multimodal seismic dataset, comprising over 16K seismic events spanning 13 years (2010 to 2023) across diverse geographical regions. Each event data integrates waveform recordings from global station networks, intensity maps, population exposure visualizations, and a comprehensive textual description within a standardized JSON format. We additionally develop MISCE, a multimodal instruction set on top of raw data to enable supervised training and evaluation of GMMs on seismic reasoning tasks ranging from basic information retrieval to complex cross modal analysis. We leverage MISCE to finetune an existing multimodal model (Unified IO 2) enhanced with a specialized timeseries encoder, which yields SeisModal, the first domain specific multimodal model for comprehensive seismic analysis. Evaluation of state of the art multimodal models on MultiSeismo reveals significant challenges, particularly with time-series data processing for general purpose models, while demonstrating SeisModal's superior performance on seismic multimodal reasoning tasks. These results prove that MultiSeismo provides a rigorous benchmark for future multimodal research in seismology and validate the success of our domain specific architectural adaptations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。