arXiv:2510.21980cs.LGmath.PR2025-10被引 1

用热力学加权图集提升适配体亲和力预测准确率

Boltzmann Graph Ensemble Embeddings for Aptamer Libraries

  • 构建基于玻尔兹曼权重的分子图集嵌入模型
  • 在有实验偏差的数据中仍能精准识别高亲和力适配体
  • 适合用于筛选低丰度但潜力大的候选适配体

生物化学中的机器学习方法通常将分子表示为成对分子间相互作用的图,用于性质与结构预测。多数方法仅基于单个图(通常是最低自由能结构)进行分析,该结构代表热力学平衡下的低能构象集合。本文提出一种热力学参数化的指数族随机图(ERGM)嵌入方法,将分子建模为玻尔兹曼加权的相互作用图集合。我们在SELEX数据集上评估该方法,这些数据因PCR扩增或测序噪声等实验偏差,常导致观测到的适配体丰度与其真实结合强度不一致,产生异常候选。结果表明,该嵌入方法在存在偏差的情况下仍能实现鲁棒的社区检测与子图级解释,有效揭示适配体-配体亲和力关系。该方法可用于筛选低丰度但具潜力的适配体候选者以进一步实验验证。

原文摘要 · Abstract (English)

Machine-learning methods in biochemistry commonly represent molecules as graphs of pairwise intermolecular interactions for property and structure predictions. Most methods operate on a single graph, typically the minimal free energy (MFE) structure, for low-energy ensembles (conformations) representative of structures at thermodynamic equilibrium. We introduce a thermodynamically parameterized exponential-family random graph (ERGM) embedding that models molecules as Boltzmann-weighted ensembles of interaction graphs. We evaluate this embedding on SELEX datasets, where experimental biases (e.g., PCR amplification or sequencing noise) can obscure true aptamer-ligand affinity, producing anomalous candidates whose observed abundance diverges from their actual binding strength. We show that the proposed embedding enables robust community detection and subgraph-level explanations for aptamer ligand affinity, even in the presence of biased observations. This approach may be used to identify low-abundance aptamer candidates for further experimental evaluation.

适配体图神经网络生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。