arXiv:2511.05728cs.LGcs.AI2025-11

用数据压缩原理发现化学基团,提升药物活性预测效果。

Compressing Chemistry Reveals Functional Groups

  • 基于最小消息长度原理,自动寻找能压缩三百万分子的子结构。
  • 发现的基团包含已知功能团及新出现的更精确功能模式。
  • 针对特定数据集定制的指纹显著优于传统指纹方法。

我们首次对传统化学功能基团在化学解释中的有效性进行大规模正式评估。评估基于计算学习理论的基本原则:好的解释应能压缩数据。我们提出一种基于最小消息长度(MML)原理的无监督学习算法,用于搜索约三百万种生物相关分子中能实现数据压缩的子结构。结果表明,所发现的子结构不仅包含大多数人工标注的功能基团,还揭示了具有更具体功能的新颖大尺度模式。我们在24个特定生物活性预测数据集上运行该算法,发现了数据集特异的功能基团。以这些基团构建的指纹在训练岭回归模型进行生物活性回归任务时,显著优于MACCS和Morgan指纹等其他指纹表示。

原文摘要 · Abstract (English)

We introduce the first formal large-scale assessment of the utility of traditional chemical functional groups as used in chemical explanations. Our assessment employs a fundamental principle from computational learning theory: a good explanation of data should also compress the data. We introduce an unsupervised learning algorithm based on the Minimum Message Length (MML) principle that searches for substructures that compress around three million biologically relevant molecules. We demonstrate that the discovered substructures contain most human-curated functional groups as well as novel larger patterns with more specific functions. We also run our algorithm on 24 specific bioactivity prediction datasets to discover dataset-specific functional groups. Fingerprints constructed from dataset-specific functional groups are shown to significantly outperform other fingerprint representations, including the MACCS and Morgan fingerprint, when training ridge regression models on bioactivity regression tasks.

化学信息学功能基团数据压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。