arXiv:2512.20924q-bio.BMcs.LG2025-12被引 3

模型靠猜作者偏好预测药物活性,而非理解化学结构。

Clever Hans in Chemistry: Chemist Style Signals Confound Activity Prediction on Public Benchmarks

  • 用分子指纹预测论文作者,准确率达60%。
  • 仅凭作者信息就能达到与结构模型相当的预测效果。
  • 适合关注数据偏见与模型可解释性的研究人员。

机器学习模型能否仅凭分子结构判断其设计者?若能,训练于文献数据的模型可能利用作者意图而非学习因果结构-活性关系。我们通过将CHEMBL实验与发表作者关联,训练了一个1,815类分类器,基于分子指纹预测作者,在骨架分割下达到60%的top-5准确率。随后,我们训练一个仅接收蛋白标识和从结构推导出的作者概率向量的活性预测模型,不直接访问分子描述符。该仅依赖作者的模型性能与有结构信息的基线相当。这揭示了一种‘聪明汉’失败模式:模型可通过推断作者目标和偏好预测生物活性,而无需真正理解化学规律。我们分析了泄露来源,提出作者无关划分方案,并建议数据集实践以分离作者意图与生物结果。

原文摘要 · Abstract (English)

Can machine learning models identify which chemist made a molecule from structure alone? If so, models trained on literature data may exploit chemist intent rather than learning causal structure-activity relationships. We test this by linking CHEMBL assays to publication authors and training a 1,815-class classifier to predict authors from molecular fingerprints, achieving 60% top-5 accuracy under scaffold-based splitting. We then train an activity model that receives only a protein identifier and an author-probability vector derived from structure, with no direct access to molecular descriptors. This author-only model achieves predictive power comparable to a simple baseline that has access to structure. This reveals a "Clever Hans" failure mode: models can predict bioactivity largely by inferring chemist goals and favorite targets without requiring a lab-independent understanding of chemistry. We analyze the sources of this leakage, propose author-disjoint splits, and recommend dataset practices to decouple chemist intent from biological outcomes.

模型偏见化学信息学可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。