用掩码自编码器预训练提升钻井数据预测精度,实证有效。
Do Masked Autoencoders Improve Downhole Prediction? An Empirical Study on Real Well Drilling Data

- 用掩码自编码器对钻井数据进行预训练,再微调预测井下参数
- 最佳模型比监督型GRU降低19.8%的预测误差,接近LSTM表现
- 宽潜空间更关键,采样率高使掩码比例影响小,适合地质钻探场景
井下钻井遥测数据存在标注不对称:地面传感器数据以1~Hz连续生成,而有标签的井下测量成本高、间断且稀少。当前机器学习方法普遍采用从零开始的全监督训练,不适应此数据特性。本文首次对掩码自编码器(MAE)在井下钻井参数预测中的预训练效果进行实证评估。基于两个公开的美国犹他州FORGE地热井数据集,共约350万时间步的多变量钻井遥测数据,系统性地在72种MAE配置中进行全因子设计实验,并与监督型LSTM和GRU基线模型在预测总泥浆体积任务上对比。结果表明,最优MAE配置相比监督型GRU基线测试均方绝对误差降低19.8%,略逊于监督型LSTM基线6.4%。设计维度分析显示,潜空间宽度是主导因素(皮尔逊相关系数r = -0.59),而掩码比例影响可忽略,这一反常现象归因于1 Hz钻井数据中高时间冗余性。研究确立了MAE预训练在钻井分析中的可行性,并明确了其最适用条件。
原文摘要 · Abstract (English)
Downhole drilling telemetry presents a fundamental labeling asymmetry: surface sensor data are generated continuously at 1~Hz, while labeled downhole measurements are costly, intermittent, and scarce. Current machine learning approaches for downhole metric prediction universally adopt fully supervised training from scratch, which is poorly suited to this data regime. We present the first empirical evaluation of masked autoencoder (MAE) pretraining for downhole drilling metric prediction. Using two publicly available Utah FORGE geothermal wells comprising approximately 3.5 million timesteps of multivariate drilling telemetry, we conduct a systematic full-factorial design space search across 72 MAE configurations and compare them against supervised LSTM and GRU baselines on the task of predicting Total Mud Volume. Results show that the best MAE configuration reduces test mean absolute error by 19.8\% relative to the supervised GRU baseline, while trailing the supervised LSTM baseline by 6.4\%. Analysis of design dimensions reveals that latent space width is the dominant architectural choice (Pearson $r = -0.59$ with test MAE), while masking ratio has negligible effect, an unexpected finding attributed to high temporal redundancy in 1~Hz drilling data. These results establish MAE pretraining as a viable paradigm for drilling analytics and identify the conditions under which it is most beneficial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。