RubikSQL通过持续学习构建知识库,提升企业级自然语言转SQL的准确率。
RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System
- 将自然语言转SQL视为持续学习任务,动态维护知识库
- 在KaggleDBQA和BIRD Mini-Dev上达到当前最优性能
- 适合需要长期演进、处理复杂术语的企业级系统研发
我们提出RubikSQL,一种新型自然语言转SQL系统,旨在应对企业级实际场景中的隐式意图和领域专有术语等关键挑战。RubikSQL将NL2SQL视为持续学习任务,需同时完成知识库维护与SQL生成。通过数据库分析、结构化信息抽取、智能规则挖掘及思维链增强的SQL分析等技术,系统逐步构建并优化其知识库。随后采用多智能体工作流,利用该知识库生成精准的SQL语句。RubikSQL在KaggleDBQA与BIRD Mini-Dev数据集上均达到当前最佳表现。最后,我们发布了RubikBench基准,一个专为捕捉工业级NL2SQL核心特征而设计的新基准,为未来研究提供重要资源。
原文摘要 · Abstract (English)
We present RubikSQL, a novel NL2SQL system designed to address key challenges in real-world enterprise-level NL2SQL, such as implicit intents and domain-specific terminology. RubikSQL frames NL2SQL as a lifelong learning task, demanding both Knowledge Base (KB) maintenance and SQL generation. RubikSQL systematically builds and refines its KB through techniques including database profiling, structured information extraction, agentic rule mining, and Chain-of-Thought (CoT)-enhanced SQL profiling. RubikSQL then employs a multi-agent workflow to leverage this curated KB, generating accurate SQLs. RubikSQL achieves SOTA performance on both the KaggleDBQA and BIRD Mini-Dev datasets. Finally, we release the RubikBench benchmark, a new benchmark specifically designed to capture vital traits of industrial NL2SQL scenarios, providing a valuable resource for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。