用低成本数据提升高维输出的少样本建模精度
Projection-based multifidelity linear regression for data-scarce applications
- 基于主成分分析降维,融合低/高保真数据进行线性回归
- 仅需12个高保真样本,中位精度提升2%-12%,R²更高
- 适合高维输出、数据稀缺的工程仿真场景
针对高维输出且训练数据昂贵的系统,本文提出两种基于投影的多保真线性回归方法,利用低成本低保真数据与少量高保真数据联合建模。通过主成分基向量实现降维,采用直接数据增强和含线性校正的数据增强策略,将多保真数据统一纳入训练集,结合保真度特异性权重的加权最小二乘法求解。引入基于邻近性的自动加权选择方案,通过交叉验证优化。在高超音速飞行器表面压力场和飞机盘式制动系统温度场的测试中,当高保真样本不超过12个时,该方法相较单保真方法在相似计算成本下实现约2%-12%的中位精度提升,并取得更高的R²值。
原文摘要 · Abstract (English)
Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. The proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2%-12% improvement in median accuracy and a higher $R^2$ score relative to single-fidelity methods at comparable computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。