arXiv:2504.07297cs.LGcond-mat.mtrl-sci2025-04被引 4

融合预训练分子嵌入,提升稀疏数据下材料属性预测精度

Data Fusion of Deep Learned Molecular Embeddings for Property Prediction

  • 用独立单任务模型的嵌入进行融合,构建新多任务模型
  • 在稀疏数据集上表现优于传统多任务模型,参数更少
  • 适合数据稀疏、属性相关性弱的材料属性预测场景

数据驱动方法如深度学习可实现高精度、高效能的材料属性预测。但在许多应用中,数据稀疏严重限制了其准确性和适用性。为提升预测性能,已有研究采用迁移学习和多任务学习等技术。然而,多任务模型的表现依赖于任务间的强相关性及数据集完整性;在稀疏且弱相关属性的数据集上,标准多任务模型表现不佳。为此,本文融合由独立预训练单任务模型生成的深度学习嵌入,构建一个继承丰富、属性特异性表征的多任务模型。通过复用(而非重新训练)这些嵌入,所提出的融合模型在性能上超越标准多任务模型,且可扩展性更强,所需可训练参数更少。我们在小分子量子化学基准数据集以及一项新整理的实验数据集(来自文献及自研量子化学与热化学计算)上验证了该方法的有效性。

原文摘要 · Abstract (English)

Data-driven approaches such as deep learning can result in predictive models for material properties with exceptional accuracy and efficiency. However, in many applications, data is sparse, severely limiting their accuracy and applicability. To improve predictions, techniques such as transfer learning and multitask learning have been used. The performance of multitask learning models depends on the strength of the underlying correlations between tasks and the completeness of the data set. Standard multitask models tend to underperform when trained on sparse data sets with weakly correlated properties. To address this gap, we fuse deep-learned embeddings generated by independent pretrained single-task models, resulting in a multitask model that inherits rich, property-specific representations. By reusing (rather than retraining) these embeddings, the resulting fused model outperforms standard multitask models and can be extended with fewer trainable parameters. We demonstrate this technique on a widely used benchmark data set of quantum chemistry data for small molecules as well as a newly compiled sparse data set of experimental data collected from literature and our own quantum chemistry and thermochemical calculations.

分子嵌入多任务学习数据稀疏属性预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。