60种科学模型在分子材料表示上高度一致,揭示了通用物理表征的潜在存在。
Universally Converging Representations of Matter Across Scientific Foundation Models
- 跨模态模型通过对比学习发现对物质的内部表征高度对齐。
- 高性能模型在训练数据范围内表征接近,但新结构下所有模型退化为低信息状态。
- 为科学大模型的泛化能力提供了可量化的评估基准,适合模型选择与蒸馏。
不同模态和架构的机器学习模型被用于预测分子、材料和蛋白质的行为。然而,它们是否学习到相似的物质内部表示仍不明确。理解其潜在结构对构建能可靠外推的科学基础模型至关重要。尽管语言和视觉领域已观察到表示收敛,科学领域的对应现象尚未系统研究。本文表明,近六十个科学模型(涵盖字符串、图、3D原子、蛋白质等模态)在广泛化学体系中,其表示高度对齐。在不同数据集上训练的模型对小分子的表示极为相似;随着机器学习势能性能提升,其表示空间也趋于收敛,表明基础模型正在学习物理现实的共同表征。我们识别出两类模型行为:在训练数据相似输入下,高性能模型表征接近,低性能模型则分散至局部次优解;而在完全不同的结构上,几乎所有模型均坍缩为低信息表示,说明当前模型仍受限于训练数据与归纳偏置,尚未编码真正通用的结构。研究确立了表示对齐作为科学模型基础泛化能力的量化基准。更广泛地,该工作可追踪模型规模增长过程中物质通用表征的涌现,并用于筛选和蒸馏跨模态、跨物质域和任务迁移效果最佳的模型。
原文摘要 · Abstract (English)
Machine learning models of vastly different modalities and architectures are being trained to predict the behavior of molecules, materials, and proteins. However, it remains unclear whether they learn similar internal representations of matter. Understanding their latent structure is essential for building scientific foundation models that generalize reliably beyond their training domains. Although representational convergence has been observed in language and vision, its counterpart in the sciences has not been systematically explored. Here, we show that representations learned by nearly sixty scientific models, spanning string-, graph-, 3D atomistic, and protein-based modalities, are highly aligned across a wide range of chemical systems. Models trained on different datasets have highly similar representations of small molecules, and machine learning interatomic potentials converge in representation space as they improve in performance, suggesting that foundation models learn a common underlying representation of physical reality. We then show two distinct regimes of scientific models: on inputs similar to those seen during training, high-performing models align closely and weak models diverge into local sub-optima in representation space; on vastly different structures from those seen during training, nearly all models collapse onto a low-information representation, indicating that today's models remain limited by training data and inductive bias and do not yet encode truly universal structure. Our findings establish representational alignment as a quantitative benchmark for foundation-level generality in scientific models. More broadly, our work can track the emergence of universal representations of matter as models scale, and for selecting and distilling models whose learned representations transfer best across modalities, domains of matter, and scientific tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。