用自动生成的高质量数据和智能强化学习,让文本转SQL更准更快。
AGRO-SQL: Agentic Group-Relative Optimization with High-Fidelity Data Synthesis
- 构建可迭代生成高精度数据的工厂,确保逻辑语义对齐
- 提出新型智能体强化学习框架,推理能力提升显著
- 适合需要精准复杂查询的数据库系统开发者
当前文本转SQL系统的发展受限于高质量训练数据稀缺及模型在复杂场景下的推理能力不足。本文提出一种双中心协同框架:数据侧构建可迭代的数据工厂,生成具备高正确性与精确语义-逻辑一致性的强化学习可用数据,并通过严格验证保障质量;模型侧引入新型智能体强化学习框架,先通过多样性感知冷启动阶段初始化稳健策略,再以群体相对策略优化(GRPO)结合环境反馈持续改进推理能力。在BIRD和Spider基准上的大量实验表明,该协同方法在单模型方法中达到最新技术水平。
原文摘要 · Abstract (English)
The advancement of Text-to-SQL systems is currently hindered by the scarcity of high-quality training data and the limited reasoning capabilities of models in complex scenarios. In this paper, we propose a holistic framework that addresses these issues through a dual-centric approach. From a Data-Centric perspective, we construct an iterative data factory that synthesizes RL-ready data characterized by high correctness and precise semantic-logic alignment, ensured by strict verification. From a Model-Centric perspective, we introduce a novel Agentic Reinforcement Learning framework. This framework employs a Diversity-Aware Cold Start stage to initialize a robust policy, followed by Group Relative Policy Optimization (GRPO) to refine the agent's reasoning via environmental feedback. Extensive experiments on BIRD and Spider benchmarks demonstrate that our synergistic approach achieves state-of-the-art performance among single-model methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。