构建首个波斯语多跳问答数据集,提升知识图谱问答准确率。
A Method for Multi-Hop Question Answering on Persian Knowledge Graph
- 基于语义分解构建5600条波斯语多跳问题数据集
- 在PeCoQ数据集上F1提升12.57%,准确率提升12.06%
- 专为波斯语知识图谱设计,适合低资源语言研究者
问答系统是信息检索技术的最新发展,能够接收自然语言中的复杂查询,并利用非结构化和结构化知识源提供准确答案。知识图谱问答(KGQA)系统通过结构化数据满足用户的信息需求,将大量事实表示为图结构。然而,尽管取得显著进展,回答多跳复杂问题仍面临挑战,尤其是在波斯语中。主要难点在于准确理解并转换多跳复杂问题为语义等价的SPARQL查询,从而从知识图谱中精确获取答案。本研究为此构建了包含5600个波斯语多跳复杂问题的数据集及其语义分解形式,并基于该数据集训练波斯语模型,提出一种基于波斯语知识图谱的复杂问题问答架构。在PeCoQ数据集上的实验表明,该方法相比最优可比方法,在F1-score上提升12.57%,准确率提升12.06%。
原文摘要 · Abstract (English)
Question answering systems are the latest evolution in information retrieval technology, designed to accept complex queries in natural language and provide accurate answers using both unstructured and structured knowledge sources. Knowledge Graph Question Answering (KGQA) systems fulfill users' information needs by utilizing structured data, representing a vast number of facts as a graph. However, despite significant advancements, major challenges persist in answering multi-hop complex questions, particularly in Persian. One of the main challenges is the accurate understanding and transformation of these multi-hop complex questions into semantically equivalent SPARQL queries, which allows for precise answer retrieval from knowledge graphs. In this study, to address this issue, a dataset of 5,600 Persian multi-hop complex questions was developed, along with their decomposed forms based on the semantic representation of the questions. Following this, Persian language models were trained using this dataset, and an architecture was proposed for answering complex questions using a Persian knowledge graph. Finally, the proposed method was evaluated against similar systems on the PeCoQ dataset. The results demonstrated the superiority of our approach, with an improvement of 12.57% in F1-score and 12.06% in accuracy compared to the best comparable method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。