arXiv:2604.22313cs.CL2026-04ACL

提出新基准框架,测试对话式数据库查询的模糊与无解问题。

CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems

论文配图:CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems
图 1 · 摘自论文原文
  • 用约束驱动生成多维度模糊查询和真实对话行为
  • 主流NL2SQL系统在复杂模糊下性能下降超过30%
  • 适合关注工业级对话系统鲁棒性的研究者

工业级对话式NL2SQL系统常面临查询模糊或无解问题,尤其在用户未充分澄清的交互场景中。现有基准通常假设单一模糊来源并依赖用户互动解决,忽略了真实失败模式。我们提出Clarity框架,可自动生成包含多维度模糊性及多样用户行为的NL2SQL基准,覆盖单轮与多轮设置。通过约束驱动管道,将可执行SQL转换为含上下文延续与模式元数据的模糊查询。在Spider和BIRD数据集上的实证评估显示,包括强大型语言模型在内的主流系统在多维模糊下性能显著下降,虽能检测模糊,但难以准确定位并解决底层模式级根源。结果表明,需提升工业级NL2SQL系统的模糊检测与消解能力。

原文摘要 · Abstract (English)

NL2SQL systems deployed in industry settings often encounter ambiguous or unanswerable queries, particularly in interactive scenarios with incomplete user clarification. Existing benchmarks typically assume a single source of ambiguity and rely on user interaction for resolution, overlooking realistic failure modes. We introduce Clarity, a framework for automatically generating an NL2SQL benchmark with multi-faceted ambiguities and diverse user behaviors across both single- and multi-turn settings. Using a constraint-driven pipeline, Clarity transforms executable SQL into ambiguous queries, augmented with grounded conversational continuations and schema-level metadata. Empirical evaluation on Spider and BIRD shows that leading NL2SQL systems, including those based on strong LLMs, suffer significant performance degradation under multi-faceted ambiguity. While these systems often detect ambiguity, they struggle to accurately localize and resolve the underlying schema-level sources. Our results highlight the need for more robust ambiguity detection and resolution in industry-grade NL2SQL systems.

NL2SQL模糊性对话系统基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。