用大模型推荐3D打印药片配方,让个性化用药更快更准
FormuLLA: A Large Language Model Approach to Generating Novel 3D Printable Formulations
- 用1400组3D打印配方微调大模型,根据药量推荐辅料
- 小模型易遗忘旧知识,参数设置影响结果稳定性
- 适合药剂研发人员,尤其关注个性化制剂设计
药物三维(3D)打印技术可实现真正个性化的给药形式。近年来,人工智能被引入以加速配方与工艺开发,显著改变传统方法。然而,现有AI工作多聚焦单一问题,未能涵盖该技术的复杂配方挑战。随着AI发展,通用智能概念兴起,系统从预测模型迈向类人推理。本研究探索在超过1400组熔融沉积成型(FDM)配方数据上微调的大语言模型(LLM),用于根据活性药物成分(API)剂量推荐合适辅料,并预测丝材机械性能。四种LLM架构被微调并系统评估了微调与生成参数配置。结果表明,Llama2在推荐辅料方面表现最佳。模型选择与参数配置显著影响性能,较小的模型存在灾难性遗忘现象。此外,我们发现:(i)即使仅有1400组数据,仍可能导致模型灾难性遗忘;(ii)标准LLM评估指标仅衡量语言能力,不反映配方可加工性;(iii)在生物医学数据上训练的模型并非总是最优。解决这些问题对推动LLM从语言能力走向可靠药物配方开发至关重要。
原文摘要 · Abstract (English)
Pharmaceutical three-dimensional (3D) printing is an advanced fabrication technology with the potential to enable truly personalised dosage forms. Recent studies have integrated artificial intelligence (AI) to accelerate formulation and process development, drastically transforming current approaches to pharmaceutical 3D printing. To date, most AI-driven efforts remain narrowly focused, while failing to account for the broader formulation challenges inherent to the technology. Recent advances in AI have introduced artificial general intelligence concepts, wherein systems extend beyond conventional predictive modelling toward more generalised, human-like reasoning. In this work, we investigate the application of large language models (LLMs), fine-tuned on a fused deposition modelling (FDM) dataset comprising over 1400 formulations, to recommend suitable excipients based on active pharmaceutical ingredient (API) dose, and predict filament mechanical properties. Four LLM architectures were fine-tuned, with systematic evaluation of both fine-tuning and generative parameter configurations. Our results demonstrate that Llama2 was best suited for recommending excipients for FDM formulations. Additionally, model selection and parameterisation significantly influence performance, with smaller LLMs exhibiting instances of catastrophic forgetting. Furthermore, we demonstrate: (i) even with relatively small dataset of over 1400 formulations, it can lead to model catastrophic forgetting; (ii) standard LLM metrics only evaluate linguistic performance but not formulation processability; and (iii) LLMs trained on biomedically-related data do not always produce the best results. Addressing these challenges is essential to advancing LLMs beyond linguistic proficiency and toward reliable systems for pharmaceutical formulation development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。