arXiv:2606.18390cs.LGq-bio.QM2026-06

MOLAR通过分离真实属性与噪声标签,提升多模态分子表征学习的鲁棒性。

MOLAR: Learning Multimodal Molecular Representations from Noisy Labels

论文配图:MOLAR: Learning Multimodal Molecular Representations from Noisy Labels
图 1 · 摘自论文原文
  • 构建双通道框架:分别建模真实属性与噪声标签的映射关系
  • 在自然噪声和可控翻转数据集上均超越主流基线模型
  • 可解释性分析揭示各模态对预测的贡献度与标签可靠性

分子属性预测中,标签噪声普遍存在——因实验测定、数据库整理或弱监督管道获取,而非直接观测的纯净生物状态。将记录标签视为可靠监督会导致模型记忆错误信息,学习误导性分子证据。在多模态分子表征学习中,图-文融合或对齐可能放大此类误差传播。为此,我们提出MOLAR:一种面向噪声标签的多模态分子表征学习框架。MOLAR将潜在真实属性推断与记录标签观测分离:图结构与文本视图分别向干净属性分布提供残差证据,而一个类别标签观测通道将该分布映射至记录标签以进行训练。此设定可推导出后验标签可靠性及模态特异性分子证据。在自然噪声分子基准与受控标签翻转基准上的实验表明,MOLAR始终优于代表性基线。可视化分析进一步验证了其可解释的可靠性与模态证据诊断能力。

原文摘要 · Abstract (English)

Motivation: Noisy labels are a common challenge in molecular property prediction because molecular annotations are often obtained from assays, curated databases, or weak annotation pipelines rather than directly observed clean biological states. Treating recorded labels as reliable supervision can cause models to memorize corrupted observations and learn misleading molecular evidence. In multimodal molecular representation learning, this issue can be amplified by graph-text fusion or alignment, which may propagate label-induced errors across modalities. Results: We propose MOLAR, a noise-aware framework for learning multimodal molecular representations from noisy labels. MOLAR separates latent clean-property inference from recorded-label observation: graph and text views contribute residual evidence to a clean-property distribution, and a categorical label-observation channel maps this distribution to recorded labels for training. This formulation derives posterior label reliability and modality-specific molecular evidence from the model. Experiments on naturally noisy molecular benchmarks and controlled label-flipping benchmarks show that MOLAR consistently outperforms representative baselines. Visualization analyses further show that MOLAR provides interpretable reliability and modality-evidence diagnostics.

分子表征噪声标签多模态学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。