用小模型和少量溶剂数据生成有效指纹,助力绿色溶剂研发。
SoDaDE: Solvent Data-Driven Embeddings with Small Transformer Models
- 基于小规模溶剂数据集与小型Transformer模型生成分子指纹。
- 在新发布数据集上预测反应产率,性能优于已有表示方法。
- 为小数据场景下的数据驱动表示提供可复用的工作流,适合绿色化学研究者。
计算表示已成为推动机器学习在化学领域发展的关键。早期依赖人工设计,如今机器学习证明可通过数据学习有意义的表示。然而化学数据集规模有限,现有表示多基于涵盖多种分子类型的宽泛数据集,缺乏对溶剂等特定体系的物理上下文信息。例如通用指纹无法体现溶剂特性。由于有害溶剂使用是化工行业主要气候问题之一,绿色溶剂替代需求激增。为此,我们提出一种新型溶剂表示方法——溶剂数据驱动嵌入(SoDaDE)。SoDaDE采用小型Transformer模型与溶剂性质数据集,生成溶剂指纹。通过在近期发布的数据集上预测反应产率,验证了其有效性,表现优于先前表示方法。本文表明,即使数据量小,也能构建高质量数据驱动指纹,并建立了一套可拓展至其他应用的工作流程。
原文摘要 · Abstract (English)
Computational representations have become crucial in unlocking the recent growth of machine learning algorithms for chemistry. Initially hand-designed, machine learning has shown that meaningful representations can be learnt from data. Chemical datasets are limited and so the representations learnt from data are generic, being trained on broad datasets which contain shallow information on many different molecule types. For example, generic fingerprints lack physical context specific to solvents. However, the use of harmful solvents is a leading climate-related issue in the chemical industry, and there is a surge of interest in green solvent replacement. To empower this research, we propose a new solvent representation scheme by developing Solvent Data Driven Embeddings (SoDaDE). SoDaDE uses a small transformer model and solvent property dataset to create a fingerprint for solvents. To showcase their effectiveness, we use SoDaDE to predict yields on a recently published dataset, outperforming previous representations. We demonstrate through this paper that data-driven fingerprints can be made with small datasets and set-up a workflow that can be explored for other applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。