用大模型挖掘肽自组装规律,加速新型生物材料发现
Learning the rules of peptide self-assembly through data mining with large language models
- 结合人工与大模型从文献中构建超千条肽自组装数据
- 机器学习模型分类准确率超80%,可高效预测组装相态
- 助力设计新肽材料,适合生物制造与药物研发人员
肽是广泛存在的生物分子,能自组装形成多种结构。尽管已有大量研究探讨其化学组成与环境因素对自组装的影响,但尚缺乏系统性整合文献数据以揭示全局规律的研究。本文通过人工校对与大语言模型辅助的文献挖掘,构建了一个肽自组装数据库,包含1000余条实验数据,涵盖肽序列、实验条件及对应组装相态。基于该数据集训练的机器学习模型在相态分类任务中表现优异,准确率超过80%。同时,针对肽文献挖掘微调的GPT模型,在信息提取上显著优于预训练模型。该流程可大幅提升筛选潜在自组装肽候选者效率,指导实验设计,深化对自组装机制的理解,为传感、催化及生物材料等应用开辟新路径。
原文摘要 · Abstract (English)
Peptides are ubiquitous and important biologically derived molecules, that have been found to self-assemble to form a wide array of structures. Extensive research has explored the impacts of both internal chemical composition and external environmental stimuli on the self-assembly behaviour of these systems. However, there is yet to be a systematic study that gathers this rich literature data and collectively examines these experimental factors to provide a global picture of the fundamental rules that govern protein self-assembly behavior. In this work, we curate a peptide assembly database through a combination of manual processing by human experts and literature mining facilitated by a large language model. As a result, we collect more than 1,000 experimental data entries with information about peptide sequence, experimental conditions and corresponding self-assembly phases. Utilizing the collected data, ML models are trained and evaluated, demonstrating excellent accuracy (>80\%) and efficiency in peptide assembly phase classification. Moreover, we fine-tune our GPT model for peptide literature mining with the developed dataset, which exhibits markedly superior performance in extracting information from academic publications relative to the pre-trained model. We find that this workflow can substantially improve efficiency when exploring potential self-assembling peptide candidates, through guiding experimental work, while also deepening our understanding of the mechanisms governing peptide self-assembly. In doing so, novel structures can be accessed for a range of applications including sensing, catalysis and biomaterials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。