改进PhaseNet在远震数据上的拾取效果,提升准确率7倍以上。
Evaluating PhaseNet on Teleseismic Data with MsPASS

- 用MsPASS构建可复现工作流,支持大规模地震数据处理与标准化训练。
- 在远震数据上训练新模型,P波拾取召回率提升741.5%,0.1秒内匹配率提高683.9%。
- 模型放大虽提准率,但推理速度大幅下降,显卡比CPU更适合大模型部署。
大量研究表明,机器学习地震波形拾取器PhaseNet在局部地震信号上表现准确,但在远震信号上性能显著下降。为解决此问题,本文提出一个可复现的MsPASS工作流,(i)实现大规模地震档案的数据准备与管理可扩展性,(ii)支持PhaseNet的标准化训练与推理。我们构建了一个包含160万条波形的对照数据集,对应美国阵列网络设施(ANF)分析师标注的远震P波拾取。该数据集证实,基于区域信号训练的PhaseNet在远震数据上表现不佳。随后,我们在ANF数据集的训练子集上从头训练PhaseNet,并在非重叠的测试子集上评估,使P波拾取召回率提升741.5%,0.1秒残差窗口内的拾取数量增加683.9%。我们还评估了不同规模模型在CPU与GPU上的表现。模型规模扩大约120倍后,精确率和召回率分别提升15.6%和23.2%。然而,大模型在NVIDIA A100 GPU上推理吞吐量下降87.2%,在128核高性能CPU节点上下降97.3%。结果表明,模型扩展在GPU上更可行,单纯增大模型并非高效提升精度的方法。
原文摘要 · Abstract (English)
Numerous studies have shown that the machine-learning picker PhaseNet produces accurate P and S picks on local earthquake signals, but its performance can degrade sharply on teleseismic signals. To address this limitation, we present a reproducible MsPASS workflow that (i) enables scalable data preparation and management for large seismic archives and (ii) supports standardized PhaseNet training and inference. We assembled a control dataset of 1.6 million waveforms linked to teleseismic P-wave picks made by analysts at the USArray Array Network Facility (ANF). The control dataset confirms that the PhaseNet model trained on regional signals performs poorly on these data. We then trained PhaseNet from scratch on the training split of the ANF control dataset and evaluated it on a non-overlapping held-out test split, increasing P-pick recall by 741.5% and yielding 683.9% more picks within a 0.1s residual window. We also evaluated PhaseNet across different model sizes on both CPUs and GPUs. Increasing the model size by about 120 times improved precision and recall by 15.6% and 23.2%, respectively. However, the scaled model reduced inference throughput by 87.2% on an NVIDIA A100 GPU and by 97.3% on a 128-core high-performance CPU node. These results indicate that scaling PhaseNet is more practical on GPUs than on CPUs, and that simply enlarging the model is not an efficient way to achieve large accuracy gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。