用元学习提升分子性质预测的准确率,同时保持模型可解释性。
Meta-Learning Linear Models for Molecular Property Prediction
- 通过元学习共享多个任务间的模型参数,实现跨任务知识迁移。
- 在不同数据集上性能比传统岭回归提升1.1到25倍。
- 适合需要高精度且可解释性的化学属性预测场景。
化学家在寻找结构-性质关系时面临高质量、一致数据集稀缺的挑战。机器学习虽显著提升了化学科学中的预测能力,但现代数据驱动方法对数据量的需求也随之增加。为应对可解释人工智能(XAI)日益增长的需求,并弥合预测精度与人类可理解性之间的差距,我们提出LAMeL——一种保留可解释性的线性元学习算法,在多个化学性质预测任务中提升了预测准确性。与传统方法将每个任务孤立处理不同,LAMeL利用元学习框架,在无共享数据的情况下识别相关任务间的共同模型参数,从而学习一个通用的功能流形,作为新未见任务的更优起点。实验表明,该方法在不同数据集上的性能相比标准岭回归提升1.1至25倍。尽管提升幅度因任务而异,但LAMeL始终优于或匹配传统线性方法,成为在准确性和可解释性均至关重要的化学性质预测中的可靠工具。
原文摘要 · Abstract (English)
Chemists in search of structure-property relationships face great challenges due to limited high quality, concordant datasets. Machine learning (ML) has significantly advanced predictive capabilities in chemical sciences, but these modern data-driven approaches have increased the demand for data. In response to the growing demand for explainable AI (XAI) and to bridge the gap between predictive accuracy and human comprehensibility, we introduce LAMeL - a Linear Algorithm for Meta-Learning that preserves interpretability while improving the prediction accuracy across multiple properties. While most approaches treat each chemical prediction task in isolation, LAMeL leverages a meta-learning framework to identify shared model parameters across related tasks, even if those tasks do not share data, allowing it to learn a common functional manifold that serves as a more informed starting point for new unseen tasks. Our method delivers performance improvements ranging from 1.1- to 25-fold over standard ridge regression, depending on the domain of the dataset. While the degree of performance enhancement varies across tasks, LAMeL consistently outperforms or matches traditional linear methods, making it a reliable tool for chemical property prediction where both accuracy and interpretability are critical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。