用正交多项式稀疏回归提升热舒适指数计算精度,误差更小且更稳定。
Approximating the universal thermal climate index using sparse regression with orthogonal polynomials
- 采用正交多项式基的稀疏回归方法优化热舒适指数计算
- 平均误差和大误差频率显著降低,精度接近理论最优
- 模型泛化能力强,仅用20%数据训练仍表现稳健
通用热气候指数(UTCI)是衡量人体对环境温度感受的生物气候指标,广泛应用于生物气候学研究并逐步成为户外热舒适度的实用标准。其计算需从多个环境参数推导,过程复杂,传统方法使用六次多项式近似,虽计算高效但误差较大。本研究提出基于正交多项式基的稀疏回归方法,实现更高精度与数值稳定性,显著降低均值误差、平均绝对误差及均方根误差,并大幅减少大误差发生频率。通过勒让德多项式基构建模型,可高效生成精度与复杂度间的帕累托前沿,系数结构稳定可解释。在仅使用20%训练数据、80%测试数据的情况下,模型仍具强泛化能力,且在自助抽样中表现一致。该方法将UTC I近似为类似傅里叶展开的正交基表达,结果在L2意义下接近理论最优。
原文摘要 · Abstract (English)
The Universal Thermal Climate Index (UTCI) is a measure of thermal comfort that quantifies how humans experience environmental conditions. Due to its robustness and versatility as a bioclimatic indicator, it has been extensively employed across a wide range of studies in bioclimatology and is increasingly used as an operational measure of outdoor thermal comfort. Calculating the UTCI value from the relevant environmental parameters is nominally not straightforward, which is why using a 6th-degree polynomial approximation has become the standard way to calculate UTCI values. Although it is computationally efficient, the error of this polynomial approximation can be substantial. The goal of this study was to develop an improved version of the polynomial approximation - one that retains comparable computational efficiency but is more robust in terms of numerical stability and substantially more accurate, particularly in reducing the frequency of larger errors. This goal was achieved using sparse orthogonal regression, namely sparse regression with an orthogonal polynomial basis, which not only substantially reduces the average errors (i.e., the mean error, the mean absolute error, and the root mean square error) but also drastically reduces the frequency of large errors. By leveraging Legendre polynomial bases, approximation models could be constructed that efficiently populate a Pareto front of accuracy versus complexity and exhibit stable, hierarchical coefficient structures across varying model capacities. Training the new approximation models over only 20% of the data, with the testing performed over the remaining 80%, highlights successful generalization, with the results being robust under bootstrapping. The decomposition effectively approximates the UTCI as a Fourier-like expansion in an orthogonal basis, yielding results near the theoretical optimum in the L2 (least squares) sense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。