用可验证方法提升钢材疲劳强度预测的可靠性,发现传统方法在高强区失效。
Distribution-Free Conformal Prediction for Steel Fatigue Strength: Marginal Validity Is Not Enough

- 引入分位数校准的无分布置信区间方法,评估不同强度区间的预测可靠性
- 高强区域覆盖率仅75.8%,远低于名义90%要求,暴露传统方法风险
- 自适应方法在保持区间宽度前提下实现全范围可靠覆盖,适合工程决策
钢构件疲劳失效的实验预测成本高昂,需在多种成分与工艺条件下测试,推动数据驱动模型研究。基于NIMS MatNavi钢疲劳数据集的研究虽报告高点预测精度,但依赖整体误差指标,难以判断个体预测可靠性及性能在疲劳强度范围内的稳定性。本文首次将符合性预测应用于钢疲劳强度,对比七种区间构建方法,在50次独立数据划分上区分边际覆盖率与特定子区域覆盖率。梯度提升点模型达到R²=0.976±0.009,平均绝对误差18.3±2.3 MPa。分割符合性预测实现边际覆盖率0.918,但在最高强度四分位区间降至0.758,该模式亦见于高斯过程基线。两种局部自适应方法改善此问题:交叉拟合标准化符合性方法在各四分位区间保持0.872–0.940覆盖率且不增加平均宽度;莫德里安组条件符合性预测获得最窄区间(0.917–0.939),但宽度增加12%,部分源于小样本下按组校准带来的更保守有限样本分位数水平。相比之下,符合化分位数回归虽恢复边际有效性,却在所有四分位区间扩大区间而未弥合条件差距。基于机器学习的疲劳强度预测若仅宣称边际覆盖率,可能掩盖关键高强区系统性不可靠性;因此,应常规评估条件覆盖率。
原文摘要 · Abstract (English)
Predicting fatigue failure in steel components experimentally is costly, requiring testing across multiple compositions and processing conditions, spurring research on data-driven prediction models. Studies using the NIMS MatNavi steel fatigue dataset often report high point-prediction accuracy but rely on aggregate error metrics, leaving uncertainty about the reliability of individual predictions and whether accuracy is consistent across the fatigue-strength spectrum. This paper is the first to apply conformal prediction to steel fatigue strength, comparing seven interval-construction methods across 50 independent data splits and distinguishing marginal coverage from coverage within specific sub-regions of the predicted property. A gradient-boosting point model achieves an R^2 of 0.976 +/- 0.009 and a mean absolute error of 18.3 +/- 2.3 MPa. Split-conformal prediction provides valid marginal coverage (0.918) but drops to 0.758 in the highest-strength quartile, where design margins are most critical, a pattern also observed with a Gaussian process baseline. Two locally-adaptive methods correct this: a cross-fitted normalised conformal method holds 0.872-0.940 across quartiles at no cost in average width, and Mondrian group-conditional conformal prediction holds the tightest band of any method (0.917-0.939) at a 12% width premium, part of which traces to the more conservative finite-sample quantile level implied by per-group calibration at this sample size. Conformalized quantile regression, by contrast, restores marginal validity but inflates intervals in every quartile without closing the conditional gap. Marginal coverage claims for ML-based fatigue-strength predictions can conceal systematic unreliability precisely where engineering decisions are most risky; therefore, conditional coverage should be routinely assessed alongside marginal coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。