arXiv:2605.05476cs.LGcs.AI2026-05

构建双用途基准,评估知识图谱构建与GNN在生物医学文本中的表现

A Unified Benchmark for Evaluating Knowledge Graph Construction Methods and Graph Neural Networks

论文配图:A Unified Benchmark for Evaluating Knowledge Graph Construction Methods and Graph Neural Networks
图 1 · 摘自论文原文
  • 用同一语料库生成两种自动构建图与一个专家标注参考图
  • 在半监督节点分类任务中验证不同图的质量对GNN性能影响
  • 提供可复现框架,支持新方法接入和对比测试

从文本自动构建的知识图谱在实际应用中日益普及,但其固有的噪声、碎片化和语义不一致会显著影响图神经网络(GNN)在下游任务中的表现。现有评估难以区分性能差异是由学习模型还是图质量导致。为此,本文提出一个双用途基准,用于联合评估:(i) GNN在噪声文本来源图上的性能;(ii) 图构建方法在下游任务中的有效性。该基准基于单一生物医学文本语料库构建,包含两种不同抽取方法生成的自动图,以及一个由专家手工构建的高质量参考图,作为性能上限。通过半监督节点分类任务,实现对图构建方法的可控比较和对GNN鲁棒性的系统评估。同时提供标准化、可复现且可扩展的评估框架,便于集成新的图抽取方法与学习模型。

原文摘要 · Abstract (English)

Knowledge graphs automatically constructed from text are increasingly used in real-world applications. However, their inherent noise, fragmentation, and semantic inconsistencies significantly affect the performance of Graph Neural Networks (GNNs) on downstream tasks. Assessing their performance and robustness remains difficult, as it is often unclear whether observed results stem from the learning model or from the quality of the constructed graph itself. In this work, we introduce a dual-purpose benchmark designed to jointly evaluate (i) the performance of GNNs on noisy, text-derived graphs and (ii) the effectiveness of graph construction methods on a downstream task. The benchmark is built in the biomedical domain from a single textual corpus and includes two automatically constructed graphs generated using different extraction methods, alongside a high-quality reference graph curated by experts that serves as an upper performance bound. This design enables controlled comparison of construction methods and systematic evaluation of GNN robustness through semi-supervised node classification. We further provide a standardized, reproducible, and extensible evaluation framework, facilitating the integration of new graph extraction methods and learning models.

知识图谱图神经网络基准测试生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。