arXiv:2512.20084cs.LGcs.AI2025-12被引 1

融合3D结构与文本描述,提升催化吸附能量预测精度。

QE-Catalytic: A Graph-Language Multimodal Base Model for Relaxed-Energy Prediction in Catalytic Adsorption

  • 用图-语言双模态融合模型联合处理原子坐标与结构文本。
  • 在OC20数据集上将吸附能预测误差降至0.486 eV,降低32%。
  • 适合需要精准能量预测或逆向设计催化剂的研究者。

吸附能是催化活性的关键描述符,其准确性取决于弛豫后总能量的预测能力。本文提出QE-Catalytic,一种将大语言模型(Qwen)与E(3)等变图Transformer(Equiformer-V2)深度融合的多模态框架,支持复杂催化表面的吸附构型性质预测与逆向设计。该模型同时利用三维原子结构与结构化文本,并通过图-文对齐将3D几何信息注入语言通道,实现无精确坐标时的高性能文本预测,还可自回归生成目标能量驱动的CIF文件用于结构设计与信息补全。在OC20数据集上,其弛豫吸附能预测的平均绝对误差从0.713 eV降至0.486 eV,显著优于CatBERTa和GAP-CATBERTa等基线模型。

原文摘要 · Abstract (English)

Adsorption energy is a key descriptor of catalytic reactivity. It is fundamentally defined as the difference between the relaxed total energy of the adsorbate-surface system and that of an appropriate reference state; therefore, the accuracy of relaxed-energy prediction directly determines the reliability of machine-learning-driven catalyst screening. E(3)-equivariant graph neural networks (GNNs) can natively operate on three-dimensional atomic coordinates under periodic boundary conditions and have demonstrated strong performance on such tasks. In contrast, language-model-based approaches, while enabling human-readable textual descriptions and reducing reliance on explicit graph -- thereby broadening applicability -- remain insufficient in both adsorption-configuration energy prediction accuracy and in distinguishing ``the same system with different configurations,'' even with graph-assisted pretraining in the style of GAP-CATBERTa. To this end, we propose QE-Catalytic, a multimodal framework that deeply couples a large language model (\textbf{Q}wen) with an E(3)-equivariant graph Transformer (\textbf{E}quiformer-V2), enabling unified support for adsorption-configuration property prediction and inverse design on complex catalytic surfaces. During prediction, QE-Catalytic jointly leverages three-dimensional structures and structured configuration text, and injects ``3D geometric information'' into the language channel via graph-text alignment, allowing it to function as a high-performance text-based predictor when precise coordinates are unavailable, while also autoregressively generating CIF files for target-energy-driven structure design and information completion. On OC20, QE-Catalytic reduces the MAE of relaxed adsorption energy from 0.713~eV to 0.486~eV, and consistently outperforms baseline models such as CatBERTa and GAP-CATBERTa across multiple evaluation protocols.

催化吸附多模态能量预测逆向设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。