arXiv:2410.16975cs.CRcs.LG2024-10被引 7

公开药物分子预测模型可能泄露训练数据隐私,尤其对小众高价值分子风险更高。

Publishing Neural Networks in Drug Discovery Might Compromise Training Data Privacy

  • 用成员推断攻击检测模型暴露的化学结构隐私风险
  • 所有数据集和模型架构均存在显著隐私漏洞,小众分子更易被攻破
  • 图结构表示与消息传递网络可降低风险,适合重视数据保密的研究者

本研究探讨了在药物发现中公开机器学习模型时,暴露机密化学结构所带来的隐私风险。我们采用成员推断攻击——一种在药物发现领域尚未充分探索的隐私评估方法——在黑盒设置下分析用于分子性质预测的神经网络。结果表明,所有测试的数据集和神经网络架构均存在显著隐私风险,且多种攻击组合会加剧风险。少数类分子(药物发现中通常最具价值)尤为脆弱。此外,将分子表示为图并使用消息传递神经网络可在一定程度上缓解风险。本文提出一个评估分类模型与分子表征隐私风险的框架。研究强调,在公开基于专有化学结构训练的神经网络前需审慎权衡,提醒机构与研究人员注意数据保密与模型开放之间的取舍。

原文摘要 · Abstract (English)

This study investigates the risks of exposing confidential chemical structures when machine learning models trained on these structures are made publicly available. We use membership inference attacks, a common method to assess privacy that is largely unexplored in the context of drug discovery, to examine neural networks for molecular property prediction in a black-box setting. Our results reveal significant privacy risks across all evaluated datasets and neural network architectures. Combining multiple attacks increases these risks. Molecules from minority classes, often the most valuable in drug discovery, are particularly vulnerable. We also found that representing molecules as graphs and using message-passing neural networks may mitigate these risks. We provide a framework to assess privacy risks of classification models and molecular representations. Our findings highlight the need for careful consideration when sharing neural networks trained on proprietary chemical structures, informing organisations and researchers about the trade-offs between data confidentiality and model openness.

隐私保护药物发现机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。