用表型驱动和证据约束,让知识图谱主动发现新假设。
A phenotype-driven and evidence-governed framework for knowledge graph enrichment and hypotheses discovery in population data
- 结合GNN与LLM,从人群数据中挖掘可解释的表型。
- 多目标优化筛选出高相关、强验证、有新意的新增关系。
- 适合生物医学研究者探索未知关联,减少大模型幻觉。
现有知识图谱构建方法多为验证已有关系,难以发现新或情境依赖的节点。本文提出一种表型驱动且证据约束的框架,实现结构化假设发现与可控图谱扩展。该方法融合图神经网络(GNN)进行表型发现、因果推断与概率推理,结合大语言模型(LLM)生成假设并提取主张,形成统一流程。框架优先选择数据支持且文献中未充分研究的关系。图谱扩展被建模为多目标优化问题,候选主张在相关性、结构验证与新颖性上联合评估。通过帕累托最优选择,识别出在确认与发现间平衡的非支配主张,避免冗余知识引入。在异构人群数据集上的实验表明,该框架能生成更具可解释性的表型,揭示情境依赖的因果结构,并产出符合数据与科学证据的高质量主张。相比基于规则和仅用LLM的基线方法,该方法在合理性、新颖性、验证度与相关性之间取得最佳平衡。在检索增强设置下,召回率@5达0.98,幻觉率降至0.05,显著提升大模型输出的可靠性。
原文摘要 · Abstract (English)
Current knowledge graph (KG) construction methods are confirmatory, focusing on recovering known relationships rather than identifying novel or context-dependent nodes. This paper proposes a phenotype-driven and evidence-governed framework that shifts the paradigm toward structured hypothesis discovery and controlled KG expansion. The approach integrates graph neural networks (GNNs) for phenotype discovery, causal inference, probabilistic reasoning and large language models (LLMs) for hypothesis generation and claim extraction within a unified pipeline. The framework prioritizes relationships that are both structurally supported by data and underexplored in the literature. KG expansion is formulated as a multi-objective optimization problem, where candidate claims are jointly evaluated in terms of relevance, structural validation and novelty. Pareto-optimal selection enables the identification of non-dominated claims that balance confirmation and discovery, avoiding trivial or redundant knowledge inclusion. Experiments on heterogeneous population datasets demonstrate that the proposed framework produces more interpretable phenotypes, reveals context-dependent causal structures and generates high-quality claims that align with both data and scientific evidence. Compared to rule-based and LLM-only baselines, the method achieves the best trade-off across plausibility, novelty, validation and relevance. In retrieval-augmented settings, it significantly improves performance (Recall@5=0.98) while reducing hallucination rates (0.05), highlighting its effectiveness in grounding LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。