arXiv:2508.01871cs.AIcs.DB2025-08被引 1

构建多轮自然语言转图查询语句数据集,解决复杂交互查询难题。

Multi-turn Natural Language to Graph Query Language Translation

  • 用大模型自动生成多轮对话式图查询数据
  • 提出MTGQL数据集,基于金融图数据库构建
  • 适合研究对话式数据库查询的学者使用

近年来,自然语言转图查询语言(NL2GQL)研究日益增多。现有方法大多聚焦于单轮转换,但实际应用中用户与图数据库的交互通常是多轮、动态且依赖上下文的。单轮方法难以应对需迭代调整、探索实体关联或追问细节的复杂场景。此外,高质量多轮NL2GQL数据集稀缺严重制约了该领域发展。为此,我们提出一种基于大语言模型(LLMs)的自动化多轮NL2GQL数据集构建方法,并据此构建了源自金融市场图数据库的MTGQL数据集,将公开发布以支持后续研究。同时,我们设计了三类基线方法,用于评估多轮NL2GQL翻译的有效性,为未来研究奠定坚实基础。

原文摘要 · Abstract (English)

In recent years, research on transforming natural language into graph query language (NL2GQL) has been increasing. Most existing methods focus on single-turn transformation from NL to GQL. In practical applications, user interactions with graph databases are typically multi-turn, dynamic, and context-dependent. While single-turn methods can handle straightforward queries, more complex scenarios often require users to iteratively adjust their queries, investigate the connections between entities, or request additional details across multiple dialogue turns. Research focused on single-turn conversion fails to effectively address multi-turn dialogues and complex context dependencies. Additionally, the scarcity of high-quality multi-turn NL2GQL datasets further hinders the progress of this field. To address this challenge, we propose an automated method for constructing multi-turn NL2GQL datasets based on Large Language Models (LLMs) , and apply this method to develop the MTGQL dataset, which is constructed from a financial market graph database and will be publicly released for future research. Moreover, we propose three types of baseline methods to assess the effectiveness of multi-turn NL2GQL translation, thereby laying a solid foundation for future research.

图数据库多轮对话数据集构建NL2GQL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。