arXiv:2507.10045cs.AIcs.CL2025-07中稿 · SEMANTiCS 2025 con…被引 1

用大模型自动翻译知识图谱查询,提升数据互通效率

Automating SPARQL Query Translations between DBpedia and Wikidata

  • 用大模型直接转换DBpedia与Wikidata的SPARQL查询
  • 从100个真实查询看,转换准确率最高达83%
  • 适合需要跨知识图谱查询的研究者和开发者

本文研究前沿大语言模型(LLM)是否能自动实现主流知识图谱(KG)模式间的SPARQL查询翻译。聚焦于DBpedia与Wikidata之间的转换,并扩展至DBLP与OpenAlex KG。通过构建两个基准测试集:第一个包含100个来自QALD-9-Plus的DBpedia-Wikidata查询对;第二个包含100个DBLP查询对,用于验证在非百科类知识图谱上的泛化能力。选取三种开源大模型(Llama-3-8B、DeepSeek-R1-Distill-Llama-70B、Mistral-Large-Instruct-2407),在零样本、少样本及两种思维链提示策略下进行测试。输出结果与标准答案对比,错误类型被系统分类。结果显示模型性能差异显著,且从Wikidata到DBpedia的翻译效果明显优于反向转换。

原文摘要 · Abstract (English)

This paper investigates whether state-of-the-art Large Language Models (LLMs) can automatically translate SPARQL between popular Knowledge Graph (KG) schemas. We focus on translations between the DBpedia and Wikidata KG, and later on DBLP and OpenAlex KG. This study addresses a notable gap in KG interoperability research by rigorously evaluating LLM performance on SPARQL-to-SPARQL translation. Two benchmarks are assembled, where the first align 100 DBpedia-Wikidata queries from QALD-9-Plus; the second contains 100 DBLP queries aligned to OpenAlex, testing generalizability beyond encyclopaedic KGs. Three open LLMs: Llama-3-8B, DeepSeek-R1-Distill-Llama-70B, and Mistral-Large-Instruct-2407 are selected based on their sizes and architectures and tested with zero-shot, few-shot, and two chain-of-thought variants. Outputs were compared with gold answers, and resulting errors were categorized. We find that the performance varies markedly across models and prompting strategies, and that translations for Wikidata to DBpedia work far better than translations for DBpedia to Wikidata.

知识图谱大模型查询翻译SPARQL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。