研究联邦学习下分子数据多样性分析,提升药物研发协作效率。
Insights into the Unknown: Federated Data Diversity Analysis on Molecular Data
- 对比三种联邦聚类方法在分布式分子数据上的表现
- 引入化学信息度量指标SF-ICF,提升评估准确性
- 揭示领域知识对联邦数据多样性分析的关键作用
人工智能正深刻影响制药药物研发,但其工业应用受限于对公开数据集的依赖,缺乏专有制药数据的规模与多样性。联邦学习(FL)为跨数据孤岛实现隐私保护下的协同建模提供了可能,但也使数据中心任务如数据多样性估计、合理数据划分及化学空间结构理解变得复杂。为此,本文研究联邦聚类方法在分离与表征分布式分子数据方面的效果。我们在八个多样化的分子数据集上,对比了三种方法:联邦kMeans(Fed-kMeans)、联邦主成分分析结合联邦kMeans(Fed-PCA+Fed-kMeans),以及联邦局部敏感哈希(Fed-LSH)与其集中式对应方法。评估采用标准数学指标和本文提出的化学信息度量指标SF-ICF。大规模基准测试与深入可解释性分析表明,融入领域知识的化学信息指标至关重要,并支持客户端层面的可解释性分析,对分子数据的联邦多样性分析具有重要意义。
原文摘要 · Abstract (English)
AI methods are increasingly shaping pharmaceutical drug discovery. However, their translation to industrial applications remains limited due to their reliance on public datasets, lacking scale and diversity of proprietary pharmaceutical data. Federated learning (FL) offers a promising approach to integrate private data into privacy-preserving, collaborative model training across data silos. This federated data access complicates important data-centric tasks such as estimating dataset diversity, performing informed data splits, and understanding the structure of the combined chemical space. To address this gap, we investigate how well federated clustering methods can disentangle and represent distributed molecular data. We benchmark three approaches, Federated kMeans (Fed-kMeans), Federated Principal Component Analysis combined with Fed-kMeans (Fed-PCA+Fed-kMeans), and Federated Locality-Sensitive Hashing (Fed-LSH), against their centralized counterparts on eight diverse molecular datasets. Our evaluation utilizes both, standard mathematical and a chemistry-informed evaluation metrics, SF-ICF, that we introduce in this work. The large-scale benchmarking combined with an in-depth explainability analysis shows the importance of incorporating domain knowledge through chemistry-informed metrics, and on-client explainability analyses for federated diversity analysis on molecular data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。