小数据下融合物理知识的机器学习,提升加工过程建模可靠性。
Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling
- 区分物理清洗与统计筛选,将筛选视为可比较的建模假设
- 单次划分下最优模型在交叉验证中排名下降6位,高斯过程表现更稳定
- 基于物理基线的残差学习提升高斯过程性能,但损害树模型
在以物理机制主导的加工过程中,实验数据量小、成本高且材料依赖性强;在此条件下,数据处理方式、评估设计及物理信息整合形式的重要性甚至超过学习算法本身。基于155组镍基718合金的喷射磨削数据集,本文提出三项方法论贡献:首先,将基于物理的数据清洗与统计数据筛选分离,并将后者视为可竞争的建模假设而非隐式预处理;其次,发现15点留出集上的模型排序极不稳定——单一划分下的胜者在10折交叉验证中排名从第1降至第7,而高斯过程(GP)变体始终占据前列;第三,系统研究不同物理集成程度,发现对紧凑物理基线进行残差学习在高斯过程中表现优异,降低方差并实现可解释分解,但会损害树基模型;贝叶斯超参数调优提升了梯度提升与支持向量回归等敏感模型,却损害多阶段混合管道模型在该样本规模下的表现。高斯过程的置信区间近似校准(名义90%水平下实际覆盖率为86%)。整体表明,在小样本昂贵工艺数据场景下,可靠模型比较需明确数据筛选假设、采用稳健评估策略,并谨慎设计物理信息融入方式。
原文摘要 · Abstract (English)
In physically dominated machining processes, experimental datasets are small, expensive, and material-specific; in this regime, data curation, evaluation design, and the form of physics integration can matter as much as the learning algorithm. Using an abrasive waterjet milling dataset ($n{=}155$, Inconel\,718), we make three methodological contributions. First, we separate physics-based data \emph{cleaning} from statistical \emph{curation} and treat the latter as competing modelling hypotheses rather than silent preprocessing. Second, we find that model rankings from a 15-point hold-out set can be unstable: the single-split winner drops from rank~1 to rank~7 under 10-fold cross-validation, while Gaussian Process (GP) variants occupy the top ranks. Third, we study a spectrum of physics integration levels and find that residual learning on a compact physics baseline is competitive for GP, yielding lower variance and an interpretable decomposition, but degrades tree-based models. Bayesian hyper parameter tuning improves parameter-sensitive baselines such as gradient boosting and SVR, yet harms multi-stage hybrid pipelines at this sample size. GP uncertainty intervals are approximately calibrated ($86\%$ empirical coverage at nominal $90\%$). The resulting picture is methodological: for small, expensive process datasets, our results suggest that, in this setting, reliable model comparison benefits from explicit curation hypotheses, robust evaluation, and careful choices about how physics enters the model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。