arXiv:2604.03304physics.chem-phcond-mat.mtrl-sci2026-04

用生成模型加速高能材料发现,解决数据少难题

Generative Chemical Language Models for Energetic Materials Discovery

论文配图:Generative Chemical Language Models for Energetic Materials Discovery
图 1 · 摘自论文原文
  • 基于预训练+微调的生成模型,扩展化学语言模型应用
  • 片段编码提升合成可行性,助力高效结构设计
  • 适合材料研发、化学生成领域研究者参考

高能材料的新发现仍面临严峻挑战,主要受限于高质量数据的稀缺。为此,我们开发了生成式分子语言模型,先在大规模化学数据上预训练,再用精心整理的高能材料数据集进行微调。该迁移学习策略将化学语言模型能力从药物领域拓展至高能材料领域,为其他数据稀疏的发现问题提供通用框架。此外,我们探讨了基于片段的分子编码对化学语言模型的优势,尤其在构建可合成结构方面表现突出。这些进展为加速满足严苛性能要求的下一代高能材料设计奠定了基础。

原文摘要 · Abstract (English)

The discovery of new energetic materials remains a pressing challenge hindered by limited availability of high-quality data. To address this, we have developed generative molecular language models that have been pretrained on extensive chemical data and then fine-tuned with curated energetic materials datasets. This transfer-learning strategy extends the chemical language model capabilities beyond the pharmacological space in which they have been predominantly developed, offering a framework applicable to other data-spare discovery problems. Furthermore, we discuss the benefits of fragment-based molecular encodings for chemical language models, in particular in constructing synthetically accessible structures. Together, these advances provide a foundation for accelerating the design of next-generation energetic materials with demanding performance requirements.

分子生成高能材料语言模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。