用大模型生成航空维修数据,设计新评测指标提升文本转SQL效果
LLM-Driven Data Generation and a Novel Soft Metric for Evaluating Text-to-SQL in Aviation MRO
- 用大模型从数据库模式生成真实问题-语句对,解决领域数据少难题
- 提出基于F1的软性评估指标,比传统准确率更细致反映结果相似度
- 在真实航空MRO数据上验证,适合工业级文本转SQL系统研发者
大语言模型(LLMs)在文本转SQL任务中的应用有望让数据访问更普及,尤其在航空维护、修理与运营(MRO)等关键行业。然而,进展受制于两大挑战:传统评估指标(如执行准确率)反馈粗糙、呈二元性,以及领域专用评估数据集稀缺。本文针对这些缺口提出解决方案。为实现更细致的评估,我们引入一种基于F1分数的“软”指标,量化生成结果与标准答案之间的信息重叠程度。为缓解数据稀缺问题,我们提出一个由大模型驱动的流水线,能从数据库模式中合成真实的问答-语句对。我们在一份真实的MRO数据库上进行了实证评估。实验表明,所提软性指标比严格准确率提供更深入的性能分析;数据生成方法在构建领域专属基准方面有效。两项贡献共同构建了一个适用于专业环境下的文本转SQL系统评估与推进的稳健框架。
原文摘要 · Abstract (English)
The application of Large Language Models (LLMs) to text-to-SQL tasks promises to democratize data access, particularly in critical industries like aviation Maintenance, Repair, and Operation (MRO). However, progress is hindered by two key challenges: the rigidity of conventional evaluation metrics such as execution accuracy, which offer coarse, binary feedback, and the scarcity of domain-specific evaluation datasets. This paper addresses these gaps. To enable more nuanced assessment, we introduce a novel F1-score-based 'soft' metric that quantifies the informational overlap between generated and ground-truth SQL results. To address data scarcity, we propose an LLM-driven pipeline that synthesizes realistic question-SQL pairs from database schemas. We demonstrate our contributions through an empirical evaluation on an authentic MRO database. Our experiments show that the proposed soft metric provides more insightful performance analysis than strict accuracy, and our data generation technique is effective in creating a domain-specific benchmark. Together, these contributions offer a robust framework for evaluating and advancing text-to-SQL systems in specialized environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。