用分层分子表示法提升聚合物机器学习,实现更准更快的性能预测与逆向设计。
Chemistry Integrated Language Model using Hierarchical Molecular Representation for Polymer Informatics
- 将化学基团编码为令牌,结合数值描述符构建分层表示框架。
- 性能预测模型推理速度提升3.5倍,准确率提高0.9–4.1个百分点。
- 支持多目标逆向设计,成功保持骨架完整性并优化负相关性质。
机器学习已推动无机物和小分子材料发现,但聚合物仍难以应用此类方法。尽管数据稀缺常被视作主要瓶颈,我们证明通过战略性分子表征可克服此限制。本文提出CI-LLM框架,融合HAPPY(聚合物重复单元的分层抽象)将化学亚结构编码为令牌,并整合数值描述符至Transformer架构中。在性能预测方面,我们的描述符增强编码器De$^3$BERTa相比基于SMILES的模型实现3.5倍更快推理速度,且在四项属性上准确率提升0.9–4.1个百分点,同时提供子结构级别的可解释性。在逆向设计方面,基于GPT的生成器可生成具备目标性能的聚合物,实现100%骨架保留,并成功优化负相关属性。该综合框架展示了正向预测与逆向设计双重能力,揭示了战略性分子表征如何推进机器学习在聚合物科学中的应用。
原文摘要 · Abstract (English)
Machine learning has transformed material discovery for inorganic compounds and small molecules, yet polymers remain largely inaccessible to these methods. While data scarcity is often cited as the primary bottleneck, we demonstrate that strategic molecular representations can overcome this limitation. We introduce CI-LLM (Chemically Informed Language Model), a framework combining HAPPY (Hierarchically Abstracted rePeat unit of PolYmer), which encodes chemical substructures as tokens, with numerical descriptors within transformer architectures. For property prediction, De$^3$BERTa, our descriptor-enriched encoder, achieves 3.5x faster inference than SMILES-based models with improved accuracy ($R^2$ score gains of 0.9-4.1 percent across four properties), while providing interpretable structure-property insights at the subgroup level. For inverse design, our GPT-based generator produces polymers with targeted properties, achieving 100 percent scaffold retention and successful multi-property optimization for negatively correlated objectives. This comprehensive framework demonstrates both forward prediction and inverse design capabilities, showcasing how strategic molecular representation advances machine learning applications in polymer science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。