CheMLFlow让科研人员一键构建可复现的化学与材料信息学全流程实验。
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications
- 模块化组件+预置流程,自动完成数据到报告的全链条任务。
- 在量子化学、物化性质等预测任务上达到文献水平性能。
- 适合需要高效复现、自动化实验的材料与药物研发研究者。
CheMLFlow 是一个开源平台,用于构建和执行科学与技术应用的端到端、高通量、智能代理式工作流。该平台针对科学机器学习开发中的常见瓶颈:研究人员需将数据获取、清洗、表征、模型训练、验证、筛选、解释和报告整合成可复现的流程,即使其核心贡献仅涉及某一个环节。CheMLFlow 提供模块化工作流组件、开箱即用的参考流程、标准化成果与评估输出,降低编排负担,支持方法与数据集间的基准对比。平台设计为可扩展、可复现且适配自动化,支持可插拔的表征与模型、确定性数据划分、明确的运行产物、批量执行及报告生成。随着科学软件向代理辅助实验发展,CheMLFlow 的配置驱动工作流与结构化输出也为编码代理提供了实用接口,使其能在人类监督下协助用户构建实验、检查结果并总结发现。本文介绍系统架构、核心工作流及基准测试,涵盖量子力学、物化性质与生物活性预测任务,性能达文献水平,并展示时间序列数据集上的应用,拓展至分子化学之外的场景。
原文摘要 · Abstract (English)
CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage. CheMLFlow provides modular workflow components, ready-to-run reference pipelines, standardized artifacts, and evaluation outputs that reduce orchestration overhead and support benchmarking across methods and datasets. The platform is designed to be extensible, reproducible, and automation friendly, with pluggable representations and models, deterministic splits, explicit run artifacts, batch execution, and report generation. As scientific software increasingly moves toward agent assisted experimentation, CheMLFlow's configuration driven workflows and structured outputs also provide a practical interface for coding agents to help users construct experiments, inspect results, and summarize findings under human supervision. This article describes the system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。