arXiv:2606.14325cs.CLcs.AI2026-06被引 1

用合成数据让小模型精准理解文本转图数据库查询

Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation

论文配图:Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
图 1 · 摘自论文原文
  • 自动生成带知识图谱上下文的合成数据用于微调
  • 小模型在多个基准上性能显著提升,逼近大模型
  • 适合需本地部署、数据主权保护的场景

属性图正被广泛用于表示异构数据源。为精确访问其信息,需要基于文本转Cypher(Text2Cypher)解析器的对话接口。本文提出一种自动合成数据生成方法,可用于微调小型大语言模型完成该任务。我们在所有主要Text2Cypher基准上进行实验,结果表明,借助该合成数据生成方法,小型LLM性能显著提升,可与更大规模的专有模型竞争。这意味着在必须本地部署的场景中,既能保障数据主权,又无需昂贵的人工标注,同时保持高精度。

原文摘要 · Abstract (English)

Property Graphs are rapidly being adopted as database frameworks for representing heterogeneous data sources. To enable precise access to the information contained in them we need conversational interfaces based on Text-To-Cypher (Text2Cypher) parsers. This paper presents an automatic synthetic data generation method that can be leveraged to fine-tune small LLMs for this task. We conduct experiments on all the major Text-To-Cypher benchmarks, demonstrating that with our synthetic data generation approach we can significantly increase the performance of small LLMs, allowing them to compete with much larger proprietary models. This means that in settings in which models must be locally deployed we can ensure data-sovereignty without sacrificing accuracy and without costly annotation campaigns.

文本转查询知识图谱小模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。