用图像看分子就能预测性质,效率比传统方法高80倍
MolSight: Molecular Property Prediction with Images

- 把分子结构转成图像,用视觉模型直接预测性质
- 5个任务排名第一,10个任务全部进入前二
- 训练效率比多模态模型低80倍,适合资源有限场景
所有已合成的分子都能以二维骨架图表示,但现代性质预测更青睐分子图、3D构象或百亿参数语言模型,带来计算与数据工程负担。我们提出首个系统性大规模视觉基分子性质预测(MPP)研究——MolSight。基于10种视觉架构、7种预训练策略和200万张分子图像,在10个下游任务上评估性能,涵盖物理性质回归、药物发现分类与量子化学预测。为应对预训练分子结构复杂度差异,我们提出一种化学感知课程学习:利用五种结构复杂度描述符将数据集划分为五层,逐步提升难度,持续优于非课程基线。结果表明,仅用一张渲染的键线图经视觉编码器处理,即可实现有竞争力的分子性质预测,即‘仅凭视觉获取化学洞察’。最优课程训练配置在5个基准中排名第一,全部10个任务均位列前二,且计算量仅为最近多模态对手的1/80。
原文摘要 · Abstract (English)
Every molecule ever synthesised can be drawn as a 2D skeletal diagram, yet in modern property prediction this universally available representation has received less focus in favour of molecular graphs, 3D conformers, or billion-parameter language models, each imposing its own computational and data-engineering overhead. We present $\textbf{MolSight}$, the first systematic large-scale study of vision-based Molecular Property Prediction (MPP). Using 10 vision architectures, 7 pre-training strategies, and $2\,M$ molecule images, we evaluate performance across 10 downstream tasks spanning physical-property regression, drug-discovery classification, and quantum-chemistry prediction. To account for the wide variation in structural complexity across pre-training molecules, we further propose a $\textbf{chemistry-informed curriculum}$: five structural complexity descriptors partition the corpus into five tiers of increasing chemical difficulty, consistently outperforming non-curriculum baselines. We show that a single rendered bond-line image, processed by a vision encoder, is sufficient for competitive molecular property prediction, i.e. $\textit{chemical insight from sight alone}$. The best curriculum-trained configuration achieves the top result on $\textbf{5 of 10}$ benchmarks and top two on $\textbf{all 10}$, at $\textbf{$\textit{80$\times$ lower}$}$ FLOPs than the nearest multi-modal competitor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。