用油井产能指数特征+预测校准,提升油气产量外推预测的准确与可信度。
Out-of-Sample Hydrocarbon Production Forecasting: Time Series Machine Learning using Productivity Index-Driven Features and Inductive Conformal Prediction
- 基于油井产能指数筛选特征,降低输入维度。
- LSTM模型在维夫与诺恩油田外推预测中误差最低(测试集MAE=19.468)。
- 采用无分布假设的校准方法,保证95%预测区间有效性,适合工业级决策。
本研究提出一种新型机器学习框架,用于提升油气产量外推预测的鲁棒性,聚焦多变量时间序列分析。方法结合源自油藏工程的产能指数(PI)驱动特征选择与归纳型校准预测(ICP),实现严格的不确定性量化。基于维夫油田(井PF14、PF12)和诺恩油田(井E1H)的历史数据,评估了LSTM、BiLSTM、GRU与XGBoost等算法对历史原油产量(OPR_H)的预测能力。所有模型均成功完成未来时段的外推预测。性能通过传统误差指标(如MAE)及预测偏差、预测方向准确率(PDA)综合评估。PI特征选择显著降低输入维度,优于传统数值模拟流程。不确定性量化采用无需分布假设的ICP框架,可确保预测区间(如95%)的严格有效性,特别适用于复杂非正态数据。结果表明,LSTM模型表现最优,在井PF14测试集上取得最低MAE(19.468),在真实外推数据上为29.638,并在诺恩井E1H上得到验证。研究证明,融合领域知识与先进机器学习技术,可显著提升油气产量预测的可靠性。
原文摘要 · Abstract (English)
This research introduces a new ML framework designed to enhance the robustness of out-of-sample hydrocarbon production forecasting, specifically addressing multivariate time series analysis. The proposed methodology integrates Productivity Index (PI)-driven feature selection, a concept derived from reservoir engineering, with Inductive Conformal Prediction (ICP) for rigorous uncertainty quantification. Utilizing historical data from the Volve (wells PF14, PF12) and Norne (well E1H) oil fields, this study investigates the efficacy of various predictive algorithms-namely Long Short-Term Memory (LSTM), Bidirectional LSTM (BiLSTM), Gated Recurrent Unit (GRU), and eXtreme Gradient Boosting (XGBoost) - in forecasting historical oil production rates (OPR_H). All the models achieved "out-of-sample" production forecasts for an upcoming future timeframe. Model performance was comprehensively evaluated using traditional error metrics (e.g., MAE) supplemented by Forecast Bias and Prediction Direction Accuracy (PDA) to assess bias and trend-capturing capabilities. The PI-based feature selection effectively reduced input dimensionality compared to conventional numerical simulation workflows. The uncertainty quantification was addressed using the ICP framework, a distribution-free approach that guarantees valid prediction intervals (e.g., 95% coverage) without reliance on distributional assumptions, offering a distinct advantage over traditional confidence intervals, particularly for complex, non-normal data. Results demonstrated the superior performance of the LSTM model, achieving the lowest MAE on test (19.468) and genuine out-of-sample forecast data (29.638) for well PF14, with subsequent validation on Norne well E1H. These findings highlight the significant potential of combining domain-specific knowledge with advanced ML techniques to improve the reliability of hydrocarbon production forecasts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。