用概率结构提升机器学习精度,减少过拟合与欠拟合。
Probabilities-Informed Machine Learning
- 将目标变量的累积分布函数等概率信息融入训练过程
- 在回归、图像去噪和分类任务中显著提升模型准确性
- 适合有历史数据或可估算概率结构的工程与科学场景
机器学习在复杂回归与分类任务中表现强劲,但其性能高度依赖训练数据质量。本文提出一种受输出函数概率结构启发的新型机器学习范式,类似物理信息机器学习,但基于概率原理而非物理定律。该方法将目标变量的概率结构(如累积分布函数)嵌入训练过程,概率信息可来自历史数据或通过结构可靠性方法在实验设计阶段估算。通过引入领域特定的概率先验,该方法提升了模型精度,有效缓解了过拟合与欠拟合风险。在回归、图像去噪和分类任务中的应用验证了其在真实世界问题上的有效性。
原文摘要 · Abstract (English)
Machine learning (ML) has emerged as a powerful tool for tackling complex regression and classification tasks, yet its success often hinges on the quality of training data. This study introduces an ML paradigm inspired by domain knowledge of the structure of output function, akin to physics-informed ML, but rooted in probabilistic principles rather than physical laws. The proposed approach integrates the probabilistic structure of the target variable (such as its cumulative distribution function) into the training process. This probabilistic information is obtained from historical data or estimated using structural reliability methods during experimental design. By embedding domain-specific probabilistic insights into the learning process, the technique enhances model accuracy and mitigates risks of overfitting and underfitting. Applications in regression, image denoising, and classification demonstrate the approach's effectiveness in addressing real-world problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。