构建原子相互作用的多模态基准,让模型学会看懂量子世界。
QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions
- 以原子对为基本单元,用10通道图像表征量子相互作用
- 能量间隙预测误差低至0.201 eV,自由能误差2.15 eV
- 适合做量子机器学习、可解释性研究与多模态模型开发
尽管分子和材料机器学习发展迅速,但多数模型仍缺乏物理可迁移性:它们仅拟合整个分子或晶体中的相关性,而非学习原子对之间的量子相互作用。然而,成键、电荷重分布、轨道杂化和电子耦合均源于这些双体相互作用,定义了多体系统中的局部量子场。我们提出QuantumCanvas,一个大规模多模态基准,将双体量子系统作为物质的基本单元。数据集涵盖2,850种元素-元素对,每对标注18个电子、热力学和几何属性,并配有十通道图像表示,源自l和m分辨的轨道密度、角场变换、共占据图及电荷密度投影。这些基于物理的图像在无显式坐标下编码空间、角度和静电对称性,为量子学习提供可解释的视觉模态。在18个目标上对八种架构进行基准测试,GATv2在能隙预测上达到0.201 eV的平均绝对误差,EGNN在HOMO和LUMO预测上分别为0.265 eV和0.274 eV。DimeNet在能量相关量上取得2.27 eV总能MAE和0.132 eV排斥能MAE,多模态融合模型实现2.15 eV Mermin自由能MAE。在QuantumCanvas上预训练可提升在QM9、MD17和CrysMTM等更大数据集上的微调收敛稳定性和泛化能力。通过统一轨道物理与基于视觉的表示学习,QuantumCanvas为通过耦合视觉与数值模态学习可迁移的量子相互作用提供了原则性且可解释的基础。数据集与模型代码已开源。
原文摘要 · Abstract (English)
Despite rapid advances in molecular and materials machine learning, most models still lack physical transferability: they fit correlations across whole molecules or crystals rather than learning the quantum interactions between atomic pairs. Yet bonding, charge redistribution, orbital hybridization, and electronic coupling all emerge from these two-body interactions that define local quantum fields in many-body systems. We introduce QuantumCanvas, a large-scale multimodal benchmark that treats two-body quantum systems as foundational units of matter. The dataset spans 2,850 element-element pairs, each annotated with 18 electronic, thermodynamic, and geometric properties and paired with ten-channel image representations derived from l- and m-resolved orbital densities, angular field transforms, co-occupancy maps, and charge-density projections. These physically grounded images encode spatial, angular, and electrostatic symmetries without explicit coordinates, providing an interpretable visual modality for quantum learning. Benchmarking eight architectures across 18 targets, we report mean absolute errors of 0.201 eV on energy gap using GATv2, 0.265 eV on HOMO and 0.274 eV on LUMO using EGNN. For energy-related quantities, DimeNet attains 2.27 eV total-energy MAE and 0.132 eV repulsive-energy MAE, while a multimodal fusion model achieves a 2.15 eV Mermin free-energy MAE. Pretraining on QuantumCanvas further improves convergence stability and generalization when fine-tuned on larger datasets such as QM9, MD17, and CrysMTM. By unifying orbital physics with vision-based representation learning, QuantumCanvas provides a principled and interpretable basis for learning transferable quantum interactions through coupled visual and numerical modalities. Dataset and model implementations are available at https://github.com/KurbanIntelligenceLab/QuantumCanvas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。