arXiv:2510.00512q-bio.MNcs.AI2025-10中稿 · ICLR

用神经符号框架融合数据与生物知识,提升基因扰动预测的准确性与可解释性。

Adaptive Data-Knowledge Alignment in Genetic Perturbation Prediction

  • 基于溯因学习构建神经符号框架,动态对齐数据与知识
  • 在多个数据集上达到最高平衡一致性指标,优于现有方法
  • 能自动发现并修正生物学知识中的错误,适合系统生物学研究者

基因扰动引起的转录响应揭示了复杂细胞系统的内在规律。尽管当前方法在预测基因扰动反应方面取得进展,但缺乏生物学可解释性,且无法系统性地更新已有知识。克服这些局限需要将数据驱动学习与已有知识进行端到端整合,但数据与知识库之间存在不一致问题,如噪声、误标注和不完整性。为此,我们提出ALIGNED(自适应知识与数据对齐框架),基于溯因学习(ABL)范式,实现神经组件与符号组件的协同对齐,并支持系统性知识精炼。我们引入平衡一致性度量,评估预测结果在数据与知识之间的双重一致性。实验表明,ALIGNED在多个基准数据集上表现最优,获得最高的平衡一致性,同时重新发现了具有生物学意义的知识。本工作推动了从被动预测向可解释、可进化机制理解的转变。

原文摘要 · Abstract (English)

The transcriptional response to genetic perturbation reveals fundamental insights into complex cellular systems. While current approaches have made progress in predicting genetic perturbation responses, they provide limited biological understanding and cannot systematically refine existing knowledge. Overcoming these limitations requires an end-to-end integration of data-driven learning and existing knowledge. However, this integration is challenging due to inconsistencies between data and knowledge bases, such as noise, misannotation, and incompleteness. To address this challenge, we propose ALIGNED (Adaptive aLignment for Inconsistent Genetic kNowledgE and Data), a neuro-symbolic framework based on the Abductive Learning (ABL) paradigm. This end-to-end framework aligns neural and symbolic components and performs systematic knowledge refinement. We introduce a balanced consistency metric to evaluate the predictions' consistency against both data and knowledge. Our results show that ALIGNED outperforms state-of-the-art methods by achieving the highest balanced consistency, while also re-discovering biologically meaningful knowledge. Our work advances beyond existing methods to enable both the transparency and the evolution of mechanistic biological understanding.

基因扰动神经符号可解释性知识精炼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。