揭示分子模型如何分离几何与组成信息,关键在任务对齐。
Information Routing in Atomistic Foundation Models: How Task Alignment and Equivariance Shape Linear Disentanglement
- 用线性投影剥离组成信息,测几何信息可访问性。
- 任务对齐使几何信息保留率提升近3倍,最高达0.533。
- 对称性通道分工明确,适合研究分子表示机制者阅读。
分子属性预测模型如何使几何与组成信息清晰分离?我们提出组合探针分解(CPD),通过线性投影移除组成信号,测量剩余几何信息对岭回归探针的可访问性。在结构异构体基准测试中,组成投影表现随机水平,而几何残差达到94.6%的成对分类准确率。在QM9数据集上,十种不同架构的模型展现线性可访问性梯度:几何信息可访问性差异达6.6倍(HOMO-LUMO间隙的R²_geom从0.081到0.533)。任务对齐起主导作用:训练于HOMO-LUMO间隙的模型(R²_geom 0.44–0.53)比能量训练模型高约0.25 R²,无论架构如何。两个独立架构的消融实验验证:PaiNN在能量重训练后从0.53降至0.31,MACE从0.44降至0.08。数据多样性部分弥补目标错位:预训练于MPTraj的MACE(R²=0.36)优于仅用QM9训练的能量模型。在MACE内部,信息按对称类型路由:L=1(矢量)通道更编码偶极矩(R²=0.59 vs. 0.38),L=0(标量)通道更编码HOMO-LUMO间隙(R²=0.76 vs. 0.34)。ViSNet中无此模式。非线性探针在残差表示上产生误导结果,在纯组成目标上恢复至R²=0.68–0.95,建议本场景使用线性探针。
原文摘要 · Abstract (English)
What determines whether a molecular property prediction model organizes its representations so that geometric and compositional information can be cleanly separated? We introduce Compositional Probe Decomposition (CPD), which linearly projects out composition signal and measures how much geometric information remains accessible to a Ridge probe. We validate CPD with four independent checks, including a structural isomer benchmark where compositional projections score at chance while geometric residuals reach 94.6\% pairwise classification accuracy. Across ten models from five architectural families on QM9, we find a \emph{linear accessibility gradient}: models differ by $6.6\times$ in geometric information accessible after composition removal ($R^2_{\mathrm{geom}}$ from 0.081 to 0.533 for HOMO-LUMO gap). Three factors explain this gradient. Task alignment dominates: models trained on HOMO-LUMO gap ($R^2_{\mathrm{geom}}$ 0.44--0.53) outscore energy-trained models by $\sim$0.25 $R^2$ regardless of architecture. Within-architecture ablations on two independent architectures confirm this: PaiNN drops from 0.53 to 0.31 when retrained on energy, and MACE drops from 0.44 to 0.08. Data diversity partially compensates for misaligned objectives, with MACE pretrained on MPTraj (0.36) outperforming QM9-only energy models. Inside MACE's representations, information routes by symmetry type: $L{=}1$ (vector) channels preferentially encode dipole moment ($R^2 = 0.59$ vs.\ 0.38 in $L{=}0$), while $L{=}0$ (scalar) channels encode HOMO-LUMO gap ($R^2 = 0.76$ vs.\ 0.34 in $L{=}1$). This pattern is absent in ViSNet. We also show that nonlinear probes produce misleading results on residualized representations, recovering $R^2 = 0.68$--$0.95$ on a purely compositional target, and recommend linear probes for this setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。