研究靶点上下文如何影响分子属性预测,发现用对方法才能有效提升性能。
When Does Context Help? A Systematic Study of Target-Conditional Molecular Property Prediction
- 用FiLM架构将靶点信息融入分子表示,效果远优于拼接或加性融合
- 数据少时上下文可救命:CYP3A4上模型AUC达0.686,单靶随机森林仅0.238
- 上下文不匹配反而拖后腿:BACE1上性能下降10.2个百分点,少样本适配不如零样本
我们首次系统研究了靶点上下文在分子属性预测中的作用,评估了10个不同蛋白家族、4种融合架构、涵盖67至9,409个训练化合物的数据场景,以及时间与随机划分的评价方式。采用基于FiLM的NestDrug架构,得出三大核心发现:第一,融合方式决定成败:FiLM比拼接高出24.2个百分点,比加性条件化高8.6个百分点;如何融合上下文比是否融合更重要。第二,上下文可实现原本不可能的预测:在数据稀缺的CYP3A4(仅67个训练化合物)上,多任务迁移学习达到0.686 AUC,而单靶随机森林仅0.238。第三,上下文可能有害:分布不匹配导致BACE1上性能下降10.2个百分点,少样本适应始终弱于零样本。此外,我们揭示标准基准存在根本缺陷:1-最近邻Tanimoto在DUD-E上无学习即达0.991 AUC,且50%活性分子泄露于训练集,使绝对指标失真。时间划分验证(训练至2020年,测试2021–2024年)稳定达到0.843 AUC,首次提供上下文条件化分子表示能泛化至未来化学空间的严谨证据。
原文摘要 · Abstract (English)
We present the first systematic study of when target context helps molecular property prediction, evaluating context conditioning across 10 diverse protein families, 4 fusion architectures, data regimes spanning 67-9,409 training compounds, and both temporal and random evaluation splits. Using NestDrug, a FiLM-based architecture that conditions molecular representations on target identity, we characterize both success and failure modes with three principal findings. First, fusion architecture dominates: FiLM outperforms concatenation by 24.2 percentage points and additive conditioning by 8.6 pp; how you incorporate context matters more than whether you include it. Second, context enables otherwise impossible predictions: on data-scarce CYP3A4 (67 training compounds), multi-task transfer achieves 0.686 AUC where per-target Random Forest collapses to 0.238. Third, context can systematically hurt: distribution mismatch causes 10.2 pp degradation on BACE1; few-shot adaptation consistently underperforms zero-shot. Beyond methodology, we expose fundamental flaws in standard benchmarking: 1-nearest-neighbor Tanimoto achieves 0.991 AUC on DUD-E without any learning, and 50% of actives leak from training data, rendering absolute performance metrics meaningless. Our temporal split evaluation (train up to 2020, test 2021-2024) achieves stable 0.843 AUC with no degradation, providing the first rigorous evidence that context-conditional molecular representations generalize to future chemical space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。