arXiv:2608.03260cs.LG2026-08

用电子密度做自监督预训练,让模型更懂分子物理特性。

ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density

论文配图:ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density
图 1 · 摘自论文原文
  • 通过扩散模型重建被破坏的电子密度场,学习可迁移表示。
  • 在六项任务上优于从零训练,小样本下提升显著。
  • 适合需要物理一致性建模的分子性质预测与生成任务。

自监督预训练在学习可迁移表征方面展现出强大潜力,但在基于电子密度的分子学习中仍鲜有探索。电子密度提供分子电子结构的连续三维描述,同时捕捉局部空间模式与全局物理量。本研究提出ED-DiT,一种面向电子密度点云的物理引导扩散变压器,通过在不同扩散噪声水平下重建被遮蔽和损坏的对数密度场来学习可复用的表征。引入电子数守恒约束以保持总电子质量。预训练编码器可适配于性质预测、开/闭壳分类、分子-电子密度检索及分子条件下的电子密度生成。在六个EDBench任务上的实验表明,ED-DiT始终优于同架构从零训练的模型,尤其在低监督条件下表现突出。在分子条件电子密度预测中,均方根误差由2.2474降至1.3753,并超越现有基线;仅使用10%标签时,轨道能预测的均方根误差从0.0293降至0.0138。结果证明了物理引导电子密度预训练在学习可迁移分子表征方面的有效性。

原文摘要 · Abstract (English)

Pretraining has shown strong potential for learning transferable representations, yet it remains underexplored for electron-density-based molecular learning. Electron density provides a continuous three-dimensional description of molecular electronic structure, capturing both local spatial patterns and global physical quantities. This raises a key question: can electron-density fields be used for self-supervised pretraining to learn a shared representation that transfers across diverse electronic-structure-related tasks? We propose ED-DiT, a physics-guided Diffusion Transformer for self-supervised pretraining on electron-density point clouds. ED-DiT learns reusable representations by reconstructing corrupted and partially masked log-density fields across diffusion noise levels. An electron-number consistency constraint is further introduced to preserve the total electronic mass. The pretrained encoder can be adapted to property prediction, open-/closed-shell classification, molecule-electron-density retrieval, and molecule-conditioned electron-density prediction. Experiments on six EDBench tasks show that ED-DiT consistently outperforms the same architecture trained from scratch, especially under limited supervision. For molecule-conditioned electron-density prediction, it reduces RMSE from 2.2474 to 1.3753 and surpasses the available baseline. With only 10% labels, it improves orbital energy prediction RMSE from 0.0293 to 0.0138. These results demonstrate the effectiveness of physics-guided electron-density pretraining for learning transferable molecular representations.

电子密度扩散模型预训练分子表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。