arXiv:2512.16545cond-mat.mtrl-scics.LG2025-12

用小数据集的机器学习精准预测铜纳米粒尺寸,突破实验优化瓶颈。

Predictive Inorganic Synthesis based on Machine Learning using Small Data sets: a case study of size-controlled Cu Nanoparticles

  • 基于25组实验数据,用集成回归模型预测纳米粒尺寸。
  • 模型准确率R²达0.74,显著优于传统统计方法(R²=0.60)。
  • 小样本下大模型无优势,经典机器学习更适实验室规模研究。

铜纳米颗粒(Cu NPs)应用广泛,但其合成对反应参数变化极为敏感,且实验优化耗时耗力,难以实现可重复和尺寸可控的合成。尽管机器学习(ML)在材料研究中前景广阔,但常受限于高质量大数据集的缺乏。本研究基于25组自研的微波辅助聚醇法合成数据,探索用机器学习预测Cu NPs尺寸。通过拉丁超立方采样高效覆盖参数空间,构建实验数据集。集成回归模型成功实现高精度预测,决定系数R²达0.74,优于经典统计方法(R²=0.60)。此外,评估了随机森林与大语言模型(LLMs)在区分大/小颗粒上的分类性能,随机森林表现中等,而LLMs在数据稀缺条件下未展现显著优势。结果表明,经精心设计的小数据集结合稳健的古典机器学习,可有效支持Cu NPs合成预测,且对实验室级研究而言,复杂模型如LLMs未必优于简单方法。

原文摘要 · Abstract (English)

Copper nanoparticles (Cu NPs) have a broad applicability, yet their synthesis is sensitive to subtle changes in reaction parameters. This sensitivity, combined with the time- and resource-intensive nature of experimental optimization, poses a major challenge in achieving reproducible and size-controlled synthesis. While Machine Learning (ML) shows promise in materials research, its application is often limited by scarcity of large high-quality experimental data sets. This study explores ML to predict the size of Cu NPs from microwave-assisted polyol synthesis using a small data set of 25 in-house performed syntheses. Latin Hypercube Sampling is used to efficiently cover the parameter space while creating the experimental data set. Ensemble regression models successfully predict particle sizes with high accuracy ($R^2 = 0.74$), outperforming classical statistical approaches ($R^2 = 0.60$). Additionally, classification models using both random forests and Large Language Models (LLMs) are evaluated to distinguish between large and small particles. While random forests show moderate performance, LLMs offer no significant advantages under data-scarce conditions. Overall, this study demonstrates that carefully curated small data sets, paired with robust classical ML, can effectively predict the synthesis of Cu NPs and highlights that for lab-scale studies, complex models like LLMs may offer limited benefit over simpler techniques.

机器学习纳米合成小样本学习铜纳米粒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。