arXiv:2509.24403cs.CLcs.DB2025-09被引 29

通过协同测试时扩展,显著提升文本转SQL的准确率。

Agentar-Scale-SQL: Advancing Text-to-SQL through Orchestrated Test-Time Scaling

  • 引入三种协同策略:内部增强、迭代优化、并行合成与淘汰选择。
  • 在BIRD基准上达到81.67%执行准确率,领先官方排行榜。
  • 框架通用性强,适配新数据库和更强语言模型。

当前最先进的文本转SQL方法在复杂基准BIRD上仍远低于人类专家表现。现有测试时扩展方法缺乏协同策略,且忽视模型内部推理过程。为此,我们提出Agentar-Scale-SQL,一种利用可扩展计算提升性能的新框架。该框架采用协同测试时扩展策略,融合三个不同视角:i) 基于强化学习增强的内在推理(内部扩展),ii) 通过迭代精炼的序列优化(顺序扩展),iii) 利用多样化生成与锦标赛选择的并行合成(并行扩展)。Agentar-Scale-SQL为通用框架,易于适配新数据库和更强大的语言模型。大量实验表明,其在BIRD基准上取得81.67%的测试集执行准确率,位居官方排行榜首位,展现出迈向人类水平性能的有效路径。

原文摘要 · Abstract (English)

State-of-the-art (SOTA) Text-to-SQL methods still lag significantly behind human experts on challenging benchmarks like BIRD. Current approaches that explore test-time scaling lack an orchestrated strategy and neglect the model's internal reasoning process. To bridge this gap, we introduce Agentar-Scale-SQL, a novel framework leveraging scalable computation to improve performance. Agentar-Scale-SQL implements an Orchestrated Test-Time Scaling strategy that synergistically combines three distinct perspectives: i) Internal Scaling via RL-enhanced Intrinsic Reasoning, ii) Sequential Scaling through Iterative Refinement, and iii) Parallel Scaling using Diverse Synthesis and Tournament Selection. Agentar-Scale-SQL is a general-purpose framework designed for easy adaptation to new databases and more powerful language models. Extensive experiments show that Agentar-Scale-SQL achieves SOTA performance on the BIRD benchmark, reaching 81.67% execution accuracy on the test set and ranking first on the official leaderboard, demonstrating an effective path toward human-level performance.

文本转SQL测试时扩展强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。