arXiv:2511.21747physics.chem-phcs.AI2025-11被引 1

构建20万分子量子化学数据集,支持光引发剂高效设计。

QuantumChem-200K: A Large-Scale Open Organic Molecular Dataset for Quantum-Chemistry Property Screening and Language Model Benchmarking

  • 融合多种计算方法构建大规模有机分子量子属性数据集
  • 基于该数据集微调大模型,在光物理性质预测上超越GPT-4o等
  • 为光敏材料发现提供首个可扩展的AI筛选平台

下一代双光子聚合光引发剂的发现受限于缺乏包含量子化学与光物理性质的大规模开放数据集。现有分子数据集通常仅提供基础理化描述符,无法支撑数据驱动筛选或AI辅助设计。为此,我们推出QuantumChem-200K,一个包含超过20万种有机分子的大型数据集,标注了11项量子化学属性,包括双光子吸收(TPA)截面、TPA光谱范围、单重-三重态系间窜越(ISC)能量、毒性与合成可及性评分、亲水性、溶解度、沸点、分子量和芳香度。这些值通过结合密度泛函理论(DFT)、半经验激发态方法、原子级量子求解器与神经网络预测器的混合工作流计算得出。利用QuantumChem-200K,我们对开源Qwen2.5-32B大语言模型进行领域微调,构建出可从SMILES进行正向属性预测的化学AI助手。在VQM24和ZINC20中3000个未见分子上的基准测试表明,领域微调显著提升精度,优于GPT-4o、Llama-3.1-70B及原始Qwen2.5-32B模型,尤其在光引发剂设计关键的TPA与ISC预测上表现突出。QuantumChem-200K与对应AI助手共同构成首个可扩展的、基于大模型的光引发剂高通量筛选平台,加速光敏材料发现。

原文摘要 · Abstract (English)

The discovery of next-generation photoinitiators for two-photon polymerization (TPP) is hindered by the absence of large, open datasets containing the quantum-chemical and photophysical properties required to model photodissociation and excited-state behavior. Existing molecular datasets typically provide only basic physicochemical descriptors and therefore cannot support data-driven screening or AI-assisted design of photoinitiators. To address this gap, we introduce QuantumChem-200K, a large-scale dataset of over 200,000 organic molecules annotated with eleven quantum-chemical properties, including two-photon absorption (TPA) cross sections, TPA spectral ranges, singlet-triplet intersystem crossing (ISC) energies, toxicity and synthetic accessibility scores, hydrophilicity, solubility, boiling point, molecular weight, and aromaticity. These values are computed using a hybrid workflow that integrates density function theory (DFT), semi-empirical excited-state methods, atomistic quantum solvers, and neural-network predictors. Using QuantumChem-200K, we fine tune the open-source Qwen2.5-32B large language model to create a chemistry AI assistant capable of forward property prediction from SMILES. Benchmarking on 3000 unseen molecules from VQM24 and ZINC20 demonstrates that domain-specific fine-tuning significantly improves accuracy over GPT-4o, Llama-3.1-70B, and the base Qwen2.5-32B model, particularly for TPA and ISC predictions central to photoinitiator design. QuantumChem-200K and the corresponding AI assistant together provide the first scalable platform for high-throughput, LLM-driven photoinitiator screening and accelerated discovery of photosensitive materials.

量子化学光引发剂大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。