arXiv:2507.00884physics.chem-phcs.AI2025-07被引 2

用线性张量注意力实现高效精准的生物分子力场建模

A Scalable and Quantum-Accurate Foundation Model for Biomolecular Force Field via Linearly Tensorized Quadrangle Attention

  • 引入TQA机制,线性复杂度建模三/四体相互作用
  • 在rMD17等数据集上超越MACE等主流模型,精度达量子力学级别
  • 支持从构象搜索到自由能面构建的全流程应用,推理速度快10倍

精确的原子级生物分子模拟对疾病机理理解、药物发现和生物材料设计至关重要,但现有方法存在显著局限。经典力场效率高但对过渡态和精细构象描述不准;量子力学方法精度高但难以用于大规模或长时间模拟。基于AI的力场(AIFF)旨在兼顾精度与效率,却受限于多体建模复杂性、训练数据不足及泛化验证缺失。为此,我们提出LiTEN,一种具备等变性的神经网络,采用张量化四边形注意力(TQA),通过向量运算重参数化高阶张量特征,以线性复杂度高效建模三/四体相互作用,避免昂贵的球谐函数计算。基于LiTEN,LiTEN-FF为预训练的AIFF基础模型,先在nablaDFT数据集上预训练以实现广泛的化学泛化,再在SPICE数据集上微调以实现溶剂体系的高精度模拟。LiTEN在rMD17、MD22和Chignolin多数评估子集上达到当前最优(SOTA)性能,优于MACE、NequIP和EquiFormer。LiTEN-FF实现了迄今最全面的下游生物分子建模任务,包括量子级构象搜索、几何优化和自由能面构建,且对大分子(约1000原子)的推理速度比MACE-OFF快10倍。本研究提出一个物理合理、高度高效的框架,推动复杂生物分子建模发展,为药物发现等应用提供通用基础。

原文摘要 · Abstract (English)

Accurate atomistic biomolecular simulations are vital for disease mechanism understanding, drug discovery, and biomaterial design, but existing simulation methods exhibit significant limitations. Classical force fields are efficient but lack accuracy for transition states and fine conformational details critical in many chemical and biological processes. Quantum Mechanics (QM) methods are highly accurate but computationally infeasible for large-scale or long-time simulations. AI-based force fields (AIFFs) aim to achieve QM-level accuracy with efficiency but struggle to balance many-body modeling complexity, accuracy, and speed, often constrained by limited training data and insufficient validation for generalizability. To overcome these challenges, we introduce LiTEN, a novel equivariant neural network with Tensorized Quadrangle Attention (TQA). TQA efficiently models three- and four-body interactions with linear complexity by reparameterizing high-order tensor features via vector operations, avoiding costly spherical harmonics. Building on LiTEN, LiTEN-FF is a robust AIFF foundation model, pre-trained on the extensive nablaDFT dataset for broad chemical generalization and fine-tuned on SPICE for accurate solvated system simulations. LiTEN achieves state-of-the-art (SOTA) performance across most evaluation subsets of rMD17, MD22, and Chignolin, outperforming leading models such as MACE, NequIP, and EquiFormer. LiTEN-FF enables the most comprehensive suite of downstream biomolecular modeling tasks to date, including QM-level conformer searches, geometry optimization, and free energy surface construction, while offering 10x faster inference than MACE-OFF for large biomolecules (~1000 atoms). In summary, we present a physically grounded, highly efficient framework that advances complex biomolecular modeling, providing a versatile foundation for drug discovery and related applications.

力场建模AI制药张量注意力分子模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。