用不确定性分解提升分子属性预测,测试时自动融合相似分子数据
Adapting Evidential Neural Networks to Test-Time Neighbor Fusion Improves Molecular Property Prediction

- 基于证据神经网络,用不确定度指导邻居分子加权融合
- 在16个数据集上平均降低19.4%的预测误差,校准性更好
- 适合需要持续更新预测的药物研发场景,无需重新训练
训练好的分子属性模型可在测试时通过融合最相似训练分子的实际标签进行优化,这一无需重训练的方法称为邻居融合;证据神经网络通过其偶然和认知不确定性,使该过程具有理论基础。本文提出PG-EVIKAL,学习一个属性-距离度量,根据属性相关性重新排序结构相似的邻居后再融合,基于EVIKAL(标量卡尔曼滤波)与GP-EVIKAL(处理相关邻居的高斯过程变体)。在16个分子数据集上,相较于证据模型基线,PG-EVIKAL在14个数据集上降低了均方根误差,中位降幅达19.4%,且提升了校准性能;在序列化实验场景中可动态纳入新测分子,随数据到达持续优化预测而无需重训练。本工作表明,证据不确定性分解不仅是校准目标,更是可操作的推理资源,支持测试时对分子属性预测的精细化修正。
原文摘要 · Abstract (English)
A trained molecular property model can be refined at test time by correcting each prediction with the measured labels of the most similar training molecules, a retraining-free procedure we call neighbor fusion; evidential neural networks make it principled by using their aleatoric and epistemic uncertainty to parameterize a Bayesian update. Our main contribution, PG-EVIKAL, learns a property-distance metric to re-rank structurally similar neighbors by their property relevance before fusion, building on EVIKAL (scalar Kalman filter) and GP-EVIKAL (Gaussian process variant handling correlated neighbors). Evaluated on 16 molecular datasets, PG-EVIKAL reduces RMSE relative to the evidential model baseline on 14 of them, with a median reduction of 19.4%, and improves calibration; in sequential-assay scenarios it further incorporates newly measured molecules, refining predictions as they arrive without retraining. This work demonstrates that evidential uncertainty decomposition is not merely a calibration objective but an actionable inference resource that enables test-time refinement of molecular property predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。