首个用于少样本学习的时间序列溶剂选择基准数据集
The Catechol Benchmark: Time-series Solvent Selection Data for Few-shot Machine Learning
- 构建覆盖1200+工艺条件的连续过程数据集
- 实现溶剂选择的少样本回归预测,支持迁移与主动学习
- 适合化工AI、可持续制造和自动化实验研究者
机器学习有望改变实验室化学格局,在分子性质预测和反应逆合成方面已取得显著成果。然而,化学数据集常因需清洗、需深入理解化学知识或无法获取而难以被机器学习社区使用。本文引入一个全新的产率预测数据集,提供首个可用于机器学习基准测试的瞬态流动数据,涵盖超过1200种工艺条件。与以往聚焦离散参数的数据集不同,本实验设置可采样大量连续过程条件,为机器学习模型带来新挑战。研究聚焦溶剂选择,这一理论建模困难、非常适合机器学习应用的任务。通过展示回归算法、迁移学习、特征工程和主动学习的基准测试,该数据集在溶剂替代和可持续制造方面具有重要应用价值。
原文摘要 · Abstract (English)
Machine learning has promised to change the landscape of laboratory chemistry, with impressive results in molecular property prediction and reaction retro-synthesis. However, chemical datasets are often inaccessible to the machine learning community as they tend to require cleaning, thorough understanding of the chemistry, or are simply not available. In this paper, we introduce a novel dataset for yield prediction, providing the first-ever transient flow dataset for machine learning benchmarking, covering over 1200 process conditions. While previous datasets focus on discrete parameters, our experimental set-up allow us to sample a large number of continuous process conditions, generating new challenges for machine learning models. We focus on solvent selection, a task that is particularly difficult to model theoretically and therefore ripe for machine learning applications. We showcase benchmarking for regression algorithms, transfer-learning approaches, feature engineering, and active learning, with important applications towards solvent replacement and sustainable manufacturing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。