用大模型生成数据,让小模型轻松搞定复杂数据库查询。
Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs
- 用大模型自动构造高质量训练数据,指导小模型学习。
- 在多个基准上准确率提升36%,客户数据上提升22.6%。
- 适合需要高效、精准数据库查询的中小企业应用。
大型语言模型(LLMs)在商业应用中的普及推动了自然语言转SQL(NL2SQL)技术的发展,但高性能与高效率之间存在矛盾,且领域和客户定制需求进一步增加挑战。为此,我们提出Distill-C,一种专为NL2SQL设计的蒸馏定制框架。该框架利用大型教师模型通过稳健可扩展的流水线生成高质量合成数据,并在此基础上微调更小、开源的模型,使其性能媲美甚至超越大一阶的教师模型。在多个挑战性基准上,Distill-C相比三个不同来源的基础模型平均执行准确率提升36%;在三个内部客户基准上,性能较基线模型提升22.6%。结果表明,Distill-C是一种高效、高性能且通用性强的轻量级NL2SQL部署方案,在保持低计算开销的同时实现优异准确性。
原文摘要 · Abstract (English)
The growing adoption of large language models (LLMs) in business applications has amplified interest in Natural Language to SQL (NL2SQL) solutions, in which there is competing demand for high performance and efficiency. Domain- and customer-specific requirements further complicate the problem. To address this conundrum, we introduce Distill-C, a distilled customization framework tailored for NL2SQL tasks. Distill-C utilizes large teacher LLMs to produce high-quality synthetic data through a robust and scalable pipeline. Finetuning smaller and open-source LLMs on this synthesized data enables them to rival or outperform teacher models an order of magnitude larger. Evaluated on multiple challenging benchmarks, Distill-C achieves an average improvement of 36% in execution accuracy compared to the base models from three distinct LLM families. Additionally, on three internal customer benchmarks, Distill-C demonstrates a 22.6% performance improvement over the base models. Our results demonstrate that Distill-C is an effective, high-performing and generalizable approach for deploying lightweight yet powerful NL2SQL models, delivering exceptional accuracies while maintaining low computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。