用随机森林预筛选特征,大幅提升小样本下SISSO的效率与精度。
Boosting SISSO Performance on Small Sample Datasets by Using Random Forests Prescreening for Complex Feature Selection
- 先用随机森林筛选关键特征,再用SISSO做符号回归。
- 45个样本时效率比原SISSO快265倍,准确率仍超0.9。
- 适合数据少但需高精度的材料科学预测任务。
在材料科学中,数据驱动方法可加速材料发现与优化,降低研发成本并提升成功率。符号回归是提取材料描述符的关键技术,其中确保独立筛选与稀疏算子(SISSO)方法尤为突出。然而,SISSO需存储完整表达空间,内存开销大,限制了其在复杂问题中的表现。为此,本文提出一种结合随机森林(RF)与SISSO的RF-SISSO算法。该方法利用随机森林进行特征预筛选,捕捉非线性关系,提升输入数据质量,进而增强回归与分类任务的准确率与效率。在包含299种材料的验证测试中,RF-SISSO在四个不同训练样本规模下均保持测试准确率高于0.9,尤其在样本量为45时,效率较原始SISSO提升265倍。鉴于实际实验中获取大规模数据成本高昂,RF-SISSO有望在有限数据条件下高效实现高精度预测,助力科学研究。
原文摘要 · Abstract (English)
In materials science, data-driven methods accelerate material discovery and optimization while reducing costs and improving success rates. Symbolic regression is a key to extracting material descriptors from large datasets, in particular the Sure Independence Screening and Sparsifying Operator (SISSO) method. While SISSO needs to store the entire expression space to impose heavy memory demands, it limits the performance in complex problems. To address this issue, we propose a RF-SISSO algorithm by combining Random Forests (RF) with SISSO. In this algorithm, the Random Forest algorithm is used for prescreening, capturing non-linear relationships and improving feature selection, which may enhance the quality of the input data and boost the accuracy and efficiency on regression and classification tasks. For a testing on the SISSO's verification problem for 299 materials, RF-SISSO demonstrates its robust performance and high accuracy. RF-SISSO can maintain the testing accuracy above 0.9 across all four training sample sizes and significantly enhancing regression efficiency, especially in training subsets with smaller sample sizes. For the training subset with 45 samples, the efficiency of RF-SISSO was 265 times higher than that of original SISSO. As collecting large datasets would be both costly and time-consuming in the practical experiments, it is thus believed that RF-SISSO may benefit scientific researches by offering a high predicting accuracy with limited data efficiently.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。