arXiv:2409.18969cs.IRcs.AI2024-09被引 1

融合查询与大模型,提升学术数据问答准确率与效率。

Integrating SPARQL and LLMs for Question Answering over Scholarly Data Sources

  • 用SPARQL查数据,分治法处理不同问题类型和来源
  • 在三个学术数据集上实现高准确率,精确匹配率达72.3%
  • 适合需要高效处理学术文献的科研人员与知识系统开发者

2024年国际语义网会议(ISWC)举办的学术链接数据问答挑战(QALD)聚焦于对多种学术数据源——包括DBLP、SemOpenAlex和基于维基百科的文本——进行问答。本文提出一种结合SPARQL查询、分治算法与预训练抽取式问答模型的方法。首先通过SPARQL获取数据,再利用分治策略应对多类型问题与多源数据,最后由模型专门处理涉及作者个人信息的问题。该方法在精确匹配(Exact Match)和F-score指标上表现良好,验证了其在学术场景下提升问答准确率与效率的潜力。

原文摘要 · Abstract (English)

The Scholarly Hybrid Question Answering over Linked Data (QALD) Challenge at the International Semantic Web Conference (ISWC) 2024 focuses on Question Answering (QA) over diverse scholarly sources: DBLP, SemOpenAlex, and Wikipedia-based texts. This paper describes a methodology that combines SPARQL queries, divide and conquer algorithms, and a pre-trained extractive question answering model. It starts with SPARQL queries to gather data, then applies divide and conquer to manage various question types and sources, and uses the model to handle personal author questions. The approach, evaluated with Exact Match and F-score metrics, shows promise for improving QA accuracy and efficiency in scholarly contexts.

知识图谱问答系统学术数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。