用大模型补全材料合成数据,提升小样本下的预测准确率。
Leveraging Large Language Models to Address Data Scarcity in Machine Learning: Applications in Graphene Synthesis
- 用大模型填补文献数据缺失,统一实验参数表述
- 使石墨烯层数分类准确率从39%提至65%(二分类)
- 适合缺乏实验数据的材料研发与小样本学习场景
材料科学中的机器学习受限于实验数据稀缺,尤其在自主实验中合成数据获取成本高、耗时长。从文献中挖掘数据则存在质量参差、格式不一、参数报告差异等问题,导致特征难以统一,连续与离散特征混合更影响小样本学习效果。本文提出利用大语言模型(LLM)增强从现有文献整理的石墨烯化学气相沉积合成小规模异构数据集的机器学习性能。策略包括通过提示工程补全缺失数据点,并利用大模型嵌入编码复杂基底命名。该方法显著提升支持向量机(SVM)在石墨烯层数分类中的表现:二分类准确率从39%升至65%,三分类准确率从52%升至72%。对比同数据上训练微调的GPT-4模型,数值分类器结合LLM增强后仍表现更优,说明单纯微调不足以应对数据稀缺问题。真正有效的方法需结合数据补全与特征空间同质化。本研究为小样本、异构数据集的机器学习提供可推广的数据增强框架。
原文摘要 · Abstract (English)
Machine learning in materials science faces challenges due to limited experimental data, as generating synthesis data is costly and time-consuming, especially with in-house experiments. Mining data from existing literature introduces issues like mixed data quality, inconsistent formats, and variations in reporting experimental parameters, complicating the creation of consistent features for the learning algorithm. Additionally, combining continuous and discrete features can hinder the learning process with limited data. Here, we propose strategies that utilize large language models (LLMs) to enhance machine learning performance on a limited, heterogeneous dataset of graphene chemical vapor deposition synthesis compiled from existing literature. These strategies include prompting modalities for imputing missing data points and leveraging large language model embeddings to encode the complex nomenclature of substrates reported in chemical vapor deposition experiments. The proposed strategies enhance graphene layer classification using a support vector machine (SVM) model, increasing binary classification accuracy from 39% to 65% and ternary accuracy from 52% to 72%. We compare the performance of the SVM and a GPT-4 model, both trained and fine-tuned on the same data. Our results demonstrate that the numerical classifier, when combined with LLM-driven data enhancements, outperforms the standalone LLM predictor, highlighting that in data-scarce scenarios, improving predictive learning with LLM strategies requires more than simple fine-tuning on datasets. Instead, it necessitates sophisticated approaches for data imputation and feature space homogenization to achieve optimal performance. The proposed strategies emphasize data enhancement techniques, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。