arXiv:2506.12231cs.LG2025-06NeurIPS被引 3

构建化学混合物性质预测基准数据集,助力药物与电池材料研发

CheMixHub: Datasets and Benchmarks for Chemical Mixture Property Prediction

  • 整合11类化学混合物任务,覆盖50万条标注数据
  • 设计多种数据划分策略,评估模型泛化能力与鲁棒性
  • 提供深度学习模型基准,适合材料与制药领域研究者使用

开发多分子体系的高效预测模型至关重要,因为几乎所有化学产品都由多种化学品组成。尽管在工业流程中极为关键,化学混合物领域在机器学习研究中仍相对未被充分探索。本文提出CheMixHub,一个涵盖11个化学混合物性质预测任务的综合性基准,数据来自7个公开数据集,共约50万条记录。该基准引入多种数据划分方法,用于评估模型在特定场景下的泛化能力和鲁棒性,为化学混合物性质预测模型的发展奠定基础。同时,我们梳理了深度学习模型在该领域的建模空间,并建立了初始基准。该数据集有望加速化学混合物的再配方、优化与发现。数据与代码可于https://github.com/chemcognition-lab/chemixhub获取。

原文摘要 · Abstract (English)

Developing improved predictive models for multi-molecular systems is crucial, as nearly every chemical product used results from a mixture of chemicals. While being a vital part of the industry pipeline, the chemical mixture space remains relatively unexplored by the Machine Learning community. In this paper, we introduce CheMixHub, a holistic benchmark for molecular mixtures, covering a corpus of 11 chemical mixtures property prediction tasks, from drug delivery formulations to battery electrolytes, totalling approximately 500k data points gathered and curated from 7 publicly available datasets. CheMixHub introduces various data splitting techniques to assess context-specific generalization and model robustness, providing a foundation for the development of predictive models for chemical mixture properties. Furthermore, we map out the modelling space of deep learning models for chemical mixtures, establishing initial benchmarks for the community. This dataset has the potential to accelerate chemical mixture development, encompassing reformulation, optimization, and discovery. The dataset and code for the benchmarks can be found at: https://github.com/chemcognition-lab/chemixhub

化学信息学混合物预测数据集机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。