构建1000+生物信息学问答对,助力知识图谱查询自动化
A large collection of bioinformatics question-query pairs over federated knowledge graphs: methodology and applications
- 基于标准格式统一整理跨机构自然语言问答与SPARQL查询
- 包含65个跨知识图谱的联邦查询,覆盖超1000个示例
- 开源可视化与智能编辑工具,适合知识图谱维护者使用
近年来,多个生命科学资源采用统一框架结构化数据,并通过相同查询语言实现互操作性。知识图谱因其通用图结构在生物信息学中广泛应用,例如yummydata.org已收录超过60个可通过SPARQL查询的知识图谱。尽管SPARQL支持强大且可跨分布图谱的查询,但对多数用户而言仍难上手。为此,许多资源提供代表性查询示例,这些示例若数量充足并以标准化机器可读格式发布,可成为机器学习的重要数据源。本文介绍了一个由瑞士生物信息研究所(SIB)多个研究团队历时数年收集的大型生物信息学问答-查询对集合,涵盖1000多个问题与对应SPARQL查询,其中65个为跨知识图谱的联邦查询。我们提出一种基于现有标准的最小元数据统一表示方法,并开发了一系列开源应用,包括查询图可视化和智能查询编辑器,可供知识图谱维护者轻松复用。我们呼吁社区采纳并扩展该方法,以丰富知识图谱元数据,提升语义网服务能力。
原文摘要 · Abstract (English)
Background. In the last decades, several life science resources have structured data using the same framework and made these accessible using the same query language to facilitate interoperability. Knowledge graphs have seen increased adoption in bioinformatics due to their advantages for representing data in a generic graph format. For example, yummydata.org catalogs more than 60 knowledge graphs accessible through SPARQL, a technical query language. Although SPARQL allows powerful, expressive queries, even across physically distributed knowledge graphs, formulating such queries is a challenge for most users. Therefore, to guide users in retrieving the relevant data, many of these resources provide representative examples. These examples can also be an important source of information for machine learning, if a sufficiently large number of examples are provided and published in a common, machine-readable and standardized format across different resources. Findings. We introduce a large collection of human-written natural language questions and their corresponding SPARQL queries over federated bioinformatics knowledge graphs (KGs) collected for several years across different research groups at the SIB Swiss Institute of Bioinformatics. The collection comprises more than 1000 example questions and queries, including 65 federated queries. We propose a methodology to uniformly represent the examples with minimal metadata, based on existing standards. Furthermore, we introduce an extensive set of open-source applications, including query graph visualizations and smart query editors, easily reusable by KG maintainers who adopt the proposed methodology. Conclusions. We encourage the community to adopt and extend the proposed methodology, towards richer KG metadata and improved Semantic Web services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。