arXiv:2511.11626physics.chem-phcond-mat.mtrl-sci2025-11被引 1

构建超大规模聚合物计算数据库,助力AI精准预测材料性能。

Omics-scale polymer computational database transferable to real-world artificial intelligence applications

  • 用自动化分子动力学模拟生成超10万种聚合物的物理性质数据。
  • 数据库越大,AI模型从仿真到现实的迁移能力越强,呈幂律增长。
  • 适合做材料研发的科研与工业界,尤其缺实验数据的场景。

构建大规模基础数据集是推动人工智能驱动科学创新的关键里程碑。然而,与自然语言处理等成熟领域相比,材料科学尤其是聚合物研究在开放数据集建设上严重滞后,主要受限于合成与性能测试成本高、化学空间庞大复杂。本研究提出PolyOmics,一个通过全自动化分子动力学模拟管道生成的组学级计算数据库,涵盖超过10⁵种聚合物材料的多种物理性质。该数据库由约260名来自48所机构的研究人员协作开发,旨在弥合学术界与产业界差距。基于PolyOmics预训练的机器学习模型可高效微调至各类实际下游任务,即使仅有少量实验数据亦能有效应用。值得注意的是,仿真到现实的迁移模型泛化能力随数据库规模增大而显著提升,呈现幂律缩放规律。这种缩放定律支持“数据越多越好”的原则,凸显超大规模计算材料数据对提升真实世界预测性能的重要性。这一前所未有的组学级数据库揭示了大量未探索的聚合物材料区域,为人工智能驱动的聚合物科学研究奠定基础。

原文摘要 · Abstract (English)

Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science, particularly polymer research, has significantly lagged in developing extensive open datasets. This lag is primarily due to the high costs of polymer synthesis and property measurements, along with the vastness and complexity of the chemical space. This study presents PolyOmics, an omics-scale computational database generated through fully automated molecular dynamics simulation pipelines that provide diverse physical properties for over $10^5$ polymeric materials. The PolyOmics database is collaboratively developed by approximately 260 researchers from 48 institutions to bridge the gap between academia and industry. Machine learning models pretrained on PolyOmics can be efficiently fine-tuned for a wide range of real-world downstream tasks, even when only limited experimental data are available. Notably, the generalisation capability of these simulation-to-real transfer models improve significantly as the size of the PolyOmics database increases, exhibiting power-law scaling. The emergence of scaling laws supports the "more is better" principle, highlighting the significance of ultralarge-scale computational materials data for improving real-world prediction performance. This unprecedented omics-scale database reveals vast unexplored regions of polymer materials, providing a foundation for AI-driven polymer science.

聚合物计算材料AI科研大数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。