arXiv:2603.17017cs.CLcs.AI2026-03

测试大模型在噪声和语言变化下的数据库查询能力

LLM NL2SQL Robustness: Surface Noise vs. Linguistic Variation in Traditional and Agentic Settings

  • 设计十类干扰测试大模型的NL2SQL能力
  • 字符错误导致性能下降,语义不变的语言变化更难处理
  • 代理模式下语言变化挑战更大,适合关注系统鲁棒性的研究者

NL2SQL系统的鲁棒性评估至关重要,因为真实数据库环境动态、嘈杂且持续变化,而传统基准通常假设静态模式和规范用户输入。本文构建包含约十类扰动的鲁棒性评估基准,在传统与代理两种设置下评估多个顶尖大语言模型(包括Grok-4.1、Gemini-3-Pro、Claude-Opus-4.6和GPT-5.2)。结果表明,这些模型在多数扰动下表现稳健;但对表面噪声(如字符级损坏)和语义不变但词法或句法改变的语言变异敏感。此外,表面噪声在传统流水线中导致更大性能下降,而语言变异在代理设置中构成更大挑战。这些发现凸显了实现鲁棒NL2SQL系统在应对语言多样性方面仍存重大困难。

原文摘要 · Abstract (English)

Robustness evaluation for Natural Language to SQL (NL2SQL) systems is essential because real-world database environments are dynamic, noisy, and continuously evolving, whereas conventional benchmark evaluations typically assume static schemas and well-formed user inputs. In this work, we introduce a robustness evaluation benchmark containing approximately ten types of perturbations and conduct evaluations under both traditional and agentic settings. We assess multiple state-of-the-art large language models (LLMs), including Grok-4.1, Gemini-3-Pro, Claude-Opus-4.6, and GPT-5.2. Our results show that these models generally maintain strong performance under several perturbations; however, notable performance degradation is observed for surface-level noise (e.g., character-level corruption) and linguistic variation that preserves semantics while altering lexical or syntactic forms. Furthermore, we observe that surface-level noise causes larger performance drops in traditional pipelines, whereas linguistic variation presents greater challenges in agentic settings. These findings highlight the remaining challenges in achieving robust NL2SQL systems, particularly in handling linguistic variability.

NL2SQL大模型鲁棒性语言变异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。