arXiv:2606.15377cs.LGcs.AI2026-06

解决地震波到时标注噪声问题,提升模型准确性

Learning Earthquake Wave Arrival Time Picking from Labels with Inaccuracies

  • 通过特征空间对齐修正错误标签,减少噪声干扰
  • 在真实微震数据上性能提升最高达28.8%
  • 无需大规模数据集,适合地质科研场景

不准确的标注数据(即标签噪声)严重威胁监督学习模型的可靠性。这种污染会误导模型建立错误的特征-标签映射,导致泛化能力下降和测试精度降低。当前地震学应用主要依赖大规模训练集或数据增强来缓解标签噪声影响,但往往耗时耗力。本文提出一种标签噪声对比鲁棒学习(LaNCoR)方法,可在无需大规模训练数据的情况下有效处理地震信号处理中的标签噪声。该方法在特征空间中对齐输入波形特征与标签表示分布,以纠正误标并减轻其对训练过程的影响。我们在两个基准模型和训练策略下,评估了LaNCoR在真实微震数据上P相到时拾取任务的表现。结果表明,LaNCoR可使各项性能指标提升最高达28.8%。该方法在地震学与地球科学建模训练中具有广阔前景。

原文摘要 · Abstract (English)

Inaccurately labeled training data, or "label noise", poses a significant threat to the integrity of supervised machine learning models. This corruption directly degrades performance by teaching the model erroneous mappings between features and labels, which leads to poor generalization and reduced accuracy on properly labeled validation and test data. Current seismological applications mainly rely on large-scale training sets or data augmentation to reduce the label-noise impact, which can be labor-intensive and costly. Here, we introduce a Label Noise-Contrastive Robust Learning (LaNCoR) approach that can effectively handle noisy labels in seismic signal processing tasks, without requiring large-scale training datasets. In this approach, the input waveform feature and label representation distributions are aligned in the feature space to correct mislabeling and reduce its impact on the training process. We present LaNCoR's performance on the task of P-phase arrival-time picking of real microseismic data using two baseline models and training approaches. Our results indicate that LaNCoR can improve performance by up to 28.8% across performance metrics. This approach holds great promise for model training in seismology and geosciences.

地震信号噪声鲁棒标签噪声机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。