用自然语言查询生物实验数据,无需编程也能准确访问
SANE Schema-aware Natural-language Evaluation of Biological Data

- 基于真实实验结构生成自动评测基准,确保查询与数据一致
- 零训练下少样本模型即可准确生成SQL,错误多因输入模糊而非代码错误
- 适合生物学家或无编程背景的研究者快速分析高通量显微数据
高通量显微技术产生大量结构化数据,记录细胞对药物扰动的响应,但访问这些数据通常需掌握SQL技能。大语言模型提供自然语言替代方案,但幻觉问题影响结果可靠性。我们提出SANE(Schema-aware Natural-language Evaluation),一种面向特定领域的文本转SQL评估新范式:基于真实实验结构、自动构建的评测基准。SANE使评估更可扩展、系统化和可复现。利用SANE,我们评估了少样本大语言模型,发现结合结构化提示与约束机制,在受限模式下无需模型训练即可实现准确查询生成。多数失败源于输入模糊或不完整,表现为过度谨慎的澄清请求或对需先消歧的问题直接作答,而非生成错误的SQL。结果表明,当配合结构感知提示时,少样本大语言模型可在明确领域内提供可靠的数据库访问。
原文摘要 · Abstract (English)
High-throughput microscopy generates large, structured datasets capturing cellular responses to pharmacological perturbations, but accessing these datasets typically requires SQL expertise. Large language models offer a natural-language alternative, yet their tendency to hallucinate raises concerns about result reliability . We present SANE Schema-Aware Natural-language Evaluation, a novel paradigm for domain-specific text-to-SQL evaluation: schema-grounded, automatically generated benchmarks tied to real and specific experimental structure. SANE makes evaluation more scalable, systematic, and reproducible. Using SANE, we evaluate a few-shot large language model and show that, under constrained schemas with structured prompting and guardrails, accurate query generation is achievable without any model training or fine-tuning. Most failures stem from ambiguous or underspecified inputs and manifest as overly cautious clarification requests or answers to queries that should first be disambiguated, rather than incorrect SQL generation. These results indicate that few-shot large language models can provide reliable database access in well-defined domains when combined with schema-aware prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。