arXiv:2410.14436q-bio.QMcs.AI2024-10被引 2

用观测数据优化生物知识图谱,提升因果推断准确性

Learning to refine domain knowledge for biological network inference

  • 结合真实数据迭代修正先验生物知识图谱
  • 在少量干预数据下准确恢复真实因果结构
  • 适合生物网络推断与知识图谱纠错场景

扰动实验能帮助生物学家发现变量间的因果关系,但数据稀疏且高维,给因果结构学习带来挑战。生物知识图谱可辅助此类推断,但因整合了多样信息,可能偏向研究充分的系统。另一种方法是通过数据模拟训练监督模型来复现合成图谱,但真实生物学模拟难度极高。本文受两者启发,提出一种基于观测数据的领域知识精炼方法。在真实与合成数据集上,该方法在有限干预数据下优于基线模型,不仅能更准确恢复真实因果图,还能识别先验知识中的错误。

原文摘要 · Abstract (English)

Perturbation experiments allow biologists to discover causal relationships between variables of interest, but the sparsity and high dimensionality of these data pose significant challenges for causal structure learning algorithms. Biological knowledge graphs can bootstrap the inference of causal structures in these situations, but since they compile vastly diverse information, they can bias predictions towards well-studied systems. Alternatively, amortized causal structure learning algorithms encode inductive biases through data simulation and train supervised models to recapitulate these synthetic graphs. However, realistically simulating biology is arguably even harder than understanding a specific system. In this work, we take inspiration from both strategies and propose an amortized algorithm for refining domain knowledge, based on data observations. On real and synthetic datasets, we show that our approach outperforms baselines in recovering ground truth causal graphs and identifying errors in the prior knowledge with limited interventional data.

因果推断生物网络知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。