arXiv:2503.22962cs.LGcond-mat.mtrl-sci2025-03被引 17

用大模型文本与分子结构融合,小数据也能精准预测高分子性能

Multimodal machine learning with large language embedding model for polymer property prediction

  • 结合Llama3文本嵌入与Uni-Mol结构嵌入,构建多模态预测模型
  • 在少量训练数据下性能媲美甚至超越需百万样本预训练的图神经网络
  • 适合材料研发人员快速筛选高性能高分子,尤其数据稀缺场景

当前大型语言模型(如GPT-4、Llama)凭借强大的计算能力和多样文本语料,在材料科学等领域展现出优异的内容理解与生成能力。为利用这些模型中蕴含的领域知识,本文提出一种简单有效的多模态架构PolyLLMem,将Llama 3生成的文本嵌入与Uni-Mol提取的分子结构嵌入相融合,用于高分子性能预测。模型中引入低秩适配(LoRA)层,基于有限的高分子数据对嵌入进行微调,提升其对聚合物SMILES表示的化学相关性。该方法在数据稀缺条件下仍能准确预测多种高分子性能,表现可比甚至优于依赖数百万样本预训练的图神经网络与基于Transformer的模型。结果表明,像Llama这样的大模型能够有效捕捉聚合物PSMILES中的化学信息,验证了多模态融合在克服数据稀缺、加速先进聚合物材料发现中的有效性。

原文摘要 · Abstract (English)

Contemporary large language models (LLMs), such as GPT-4 and Llama, have harnessed extensive computational power and diverse text corpora to achieve remarkable proficiency in interpreting and generating domain-specific content, including materials science. To leverage the domain knowledge embedded within these models, we propose a simple yet effective multimodal architecture, PolyLLMem, which integrates text embeddings generated by Llama 3 with molecular structure embeddings derived from Uni-Mol, for polymer properties prediction tasks. In our model, Low-rank adaptation (LoRA) layers were also incorporated during the property prediction tasks to refine the embeddings based on our limited polymer dataset, thereby enhancing their chemical relevance for polymer SMILES representation. This balanced fusion of fine-tuned textual and structural information enables PolyLLMem to accurately predict a variety of polymer properties despite the scarcity of training data. Its performance is comparable to, and in some cases exceeds, that of graph-based models, as well as transformer-based models that typically require pretraining on millions of polymer samples. These findings demonstrate that LLM, such as Llama, can effectively capture chemical information encoded in polymer PSMILES, and underscore the efficacy of multimodal fusion of LLM embeddings and molecular structure embeddings in overcoming data scarcity and accelerating the discovery of advanced polymeric materials.

高分子预测多模态融合大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。