用开放数据与云端计算构建可协作的数字孪生实验室,加速科学优化。
A collaborative digital twin built on FAIR data and compute infrastructure
- 基于FAIR数据和nanoHUB平台,实现分布式实验数据共享与自动处理。
- 通过主动学习动态更新模型,指导实验设计,高效找到最优染料配方。
- 适合学生、研究者快速搭建低成本实验协作系统,通用性强。
将机器学习与自动化实验结合的自动驾驶实验室(SDL)能显著加速科学与工程中的发现与优化任务。当依托可查找、可访问、可互操作、可重用(FAIR)的数据基础设施时,具有共同兴趣的SDL可更高效协作。本文提出一种基于nanoHUB服务的分布式SDL实现,支持在线仿真与FAIR数据管理。地理分散的研究者在独立开展优化任务时,将原始实验数据提交至共享中心数据库,随后可使用自动更新的分析工具与机器学习模型。新数据通过简单网页界面提交,由nanoHUB Sim2L自动提取衍生量,并将所有输入输出索引至名为ResultsDB的FAIR数据仓库。另一套nanoHUB工作流支持基于主动学习的序列优化:研究者定义目标后,模型实时训练并指导未来实验选择。受“节俭孪生”理念启发,优化任务聚焦于通过组合食品染料实现目标颜色。利用易得且廉价的材料,研究人员和学生可自行搭建实验,共享数据,并探索FAIR数据、预测模型与序列优化的结合。所提工具通用性强,易于扩展至其他优化问题。
原文摘要 · Abstract (English)
The integration of machine learning with automated experimentation in self-driving laboratories (SDL) offers a powerful approach to accelerate discovery and optimization tasks in science and engineering applications. When supported by findable, accessible, interoperable, and reusable (FAIR) data infrastructure, SDLs with overlapping interests can collaborate more effectively. This work presents a distributed SDL implementation built on nanoHUB services for online simulation and FAIR data management. In this framework, geographically dispersed collaborators conducting independent optimization tasks contribute raw experimental data to a shared central database. These researchers can then benefit from analysis tools and machine learning models that automatically update as additional data become available. New data points are submitted through a simple web interface and automatically processed using a nanoHUB Sim2L, which extracts derived quantities and indexes all inputs and outputs in a FAIR data repository called ResultsDB. A separate nanoHUB workflow enables sequential optimization using active learning, where researchers define the optimization objective, and machine learning models are trained on-the-fly with all existing data, guiding the selection of future experiments. Inspired by the concept of ``frugal twin", the optimization task seeks to find the optimal recipe to combine food dyes to achieve the desired target color. With easily accessible and inexpensive materials, researchers and students can set up their own experiments, share data with collaborators, and explore the combination of FAIR data, predictive ML models, and sequential optimization. The tools introduced are generally applicable and can easily be extended to other optimization problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。