用大模型理解微流控表格数据,提升设计预测精度。
Autonomous Droplet Microfluidic Design Framework with Large Language Models
- 将表格数据转为文本,用大模型捕捉上下文信息
- 滴径与生成率误差降5-7倍,分类准确率提升4%以上
- 适合微流控自动化设计与生物实验研究者
基于液滴的微流控装置在生物研究中具有成本低的优势。现有机器学习模型依赖表格数据(如设计参数与效率输出),但忽略列名和描述等上下文信息。本研究提出MicroFluidic-LLMs框架,通过将表格内容转化为语言形式,利用预训练大语言模型(LLMs)进行特征提取与分析。我们在11个预测任务上评估该框架,涵盖几何、流态、工作区间和性能等维度,使用公开的流动聚焦液滴微流控数据集。结果表明,该框架可显著提升深度神经网络表现,减少繁琐的数据预处理。结合DistilBERT和GPT-2等先进自然语言模型时,滴径预测的均方绝对误差降低近5倍,生成率误差降低近7倍,工作区分类准确率提升超4%,优于此前研究。本工作为大模型在更广泛微流控应用中的潜力奠定基础。
原文摘要 · Abstract (English)
Droplet-based microfluidic devices have substantial promise as cost-effective alternatives to current assessment tools in biological research. Moreover, machine learning models that leverage tabular data, including input design parameters and their corresponding efficiency outputs, are increasingly utilised to automate the design process of these devices and to predict their performance. However, these models fail to fully leverage the data presented in the tables, neglecting crucial contextual information, including column headings and their associated descriptions. This study presents MicroFluidic-LLMs, a framework designed for processing and feature extraction, which effectively captures contextual information from tabular data formats. MicroFluidic-LLMs overcomes processing challenges by transforming the content into a linguistic format and leveraging pre-trained large language models (LLMs) for analysis. We evaluate our MicroFluidic-LLMs framework on 11 prediction tasks, covering aspects such as geometry, flow conditions, regimes, and performance, utilising a publicly available dataset on flow-focusing droplet microfluidics. We demonstrate that our MicroFluidic-LLMs framework can empower deep neural network models to be highly effective and straightforward while minimising the need for extensive data preprocessing. Moreover, the exceptional performance of deep neural network models, particularly when combined with advanced natural language processing models such as DistilBERT and GPT-2, reduces the mean absolute error in the droplet diameter and generation rate by nearly 5- and 7-fold, respectively, and enhances the regime classification accuracy by over 4%, compared with the performance reported in a previous study. This study lays the foundation for the huge potential applications of LLMs and machine learning in a wider spectrum of microfluidic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。