arXiv:2603.28325cs.CEcs.AI2026-03

用大模型从全文文献中提取疾病相关证据,构建可追溯的结构化知识图谱。

Building evidence-based knowledge bases from full-text literature for disease-specific biomedical reasoning

  • 用大模型自动提取实验性发现,生成带证据质量评分的结构化记录。
  • 发布肝癌和结直肠癌两个数据集,分别含7872、6622条证据记录。
  • 支持问答与靶点预测等下游任务,适合精准医学研究者使用。

生物医学知识资源通常要么保留未结构化的文本证据,要么压缩为缺乏研究设计、来源和量化支持的扁平三元组。本文提出EvidenceNet,一个基于全文文献的疾病特异性证据集合与图表示数据集。EvidenceNet采用大语言模型辅助流程,提取实验依据的发现并转化为结构化证据记录,对生物实体进行归一化,评估证据质量,并通过类型化语义关系连接相关记录。我们发布了EvidenceNet-HCC(7,872条证据记录,10,328个节点,49,756条边)和EvidenceNet-CRC(6,622条记录,8,795个节点,39,361条边)。技术验证显示高组件保真度:字段级提取准确率达98.3%,高置信度实体链接准确率100.0%,融合完整性87.5%,语义关系类型准确率90.0%。下游分析表明该数据支持检索增强型问答及基于图的任务,如未来链接预测与靶点优先排序。这些结果确立EvidenceNet作为面向证据感知分析与复用的疾病特异性生物医学知识库数据集。

原文摘要 · Abstract (English)

Biomedical knowledge resources often either preserve evidence as unstructured text or compress it into flat triples that omit study design, provenance, and quantitative support. Here we present EvidenceNet, a disease-specific dataset of record-level evidence collections and corresponding graph representations derived from full-text biomedical literature. EvidenceNet uses a large language model (LLM)-assisted pipeline to extract experimentally grounded findings as structured evidence records, normalize biomedical entities, score evidence quality, and connect related records through typed semantic relations. We release EvidenceNet-HCC with 7,872 evidence records and a corresponding graph with 10,328 nodes and 49,756 edges, and EvidenceNet-CRC with 6,622 records and a corresponding graph with 8,795 nodes and 39,361 edges. Technical validation shows high component fidelity, including 98.3% field-level extraction accuracy, 100.0% high-confidence entity-link accuracy, 87.5% fusion integrity, and 90.0% semantic relation-type accuracy. Downstream analyses show that the data support retrieval-augmented question answering and graph-based tasks such as future link prediction and target prioritization. These results establish EvidenceNet as a disease-specific biomedical knowledge base dataset for evidence-aware analysis and reuse.

知识图谱生物医学证据抽取大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。