让大模型读论文就能预测聚合物性能,不靠化学结构只靠文字描述。
Can LLMs Predict Polymer Physics Just by Reading Synthesis and Processing Prose?

- 用大模型直接读文献中的合成与加工描述,跳过传统化学结构输入。
- 在22项性能上中位R²达0.74,部分关键属性预测准确率超0.80。
- 适合材料研发、文献挖掘和无需结构数据的性能预测场景。
大语言模型能否仅通过阅读非结构化的科学文本预测聚合物的物理与力学性能?聚合物性能不仅由化学结构决定,更受合成路径、加工历史、形貌及测试条件影响。但现有模型多依赖结构表示(如SMILES或分子图),忽略实验上下文。本文提出PolyLM,一种完全基于自然语言、具备过程与条件感知能力的框架,直接从全文文献预测材料性能。为训练该框架,我们构建了涵盖18.5万篇论文、超过27.6万种独特聚合物样本的文献级数据集,覆盖22种物理、机械和热学性能。采用90亿参数的Qwen3.5-9B模型,结合低秩适配(LoRA)与任务级不确定性加权进行微调。在68,283个保留观测上评估,模型实现显著高精度,建立复杂性能预测新基准。在22个目标中,中位R²达0.74,关键热学、力学和理化性能预测常超越R² 0.80。结果明确表明,自然语言是真实材料性能预测的强大、可扩展接口。
原文摘要 · Abstract (English)
Can large language models predict physical and mechanical polymer properties simply by reading unstructured scientific prose? Polymer performance is rarely determined by chemical structure alone; identical nominal polymers can exhibit drastically different behaviors depending on their synthesis route, processing history, morphology, and testing conditions. Yet, state-of-the-art polymer property models typically rely on structure-only representations -- such as SMILES or molecular graphs -- which strip away this vital experimental context. In this work, we introduce \textbf{PolyLM}, a natural-language-only, process- and condition-aware framework that predicts materials performance directly from full-text literature. By circumventing structural inputs entirely, PolyLM preserves the nuanced, unstructured descriptions of synthesis and processing reported by domain scientists. To train this framework, we curated an unprecedented, literature-scale dataset encompassing 185,000 scientific papers and over 276,400 unique polymer samples across 22 physical, mechanical, and thermal properties. We fine-tuned a massive 9-billion-parameter language model (Qwen3.5-9B) using Low-Rank Adaptation (LoRA) and task-level uncertainty weighting. Evaluated on 68,283 held-out observations, the model achieves remarkably high predictive accuracy, establishing new state-of-the-art benchmarks for complex properties. Across the 22 diverse targets, the model achieves a median $R^2$ of 0.74, with predictions for key thermal, mechanical, and physicochemical properties frequently surpassing an $R^2$ of 0.80. These results unequivocally demonstrate that natural language is a powerful, highly scalable interface for realistic materials performance prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。