首个面向大模型的数据库问答基准,覆盖中英文20万+问题对。
Revolutionizing Database Q&A with Large Language Models: Comprehensive Benchmark and Evaluation
- 用大模型自动生成清洗重构数据集,覆盖手册、社区与实例知识。
- 评测9个大模型问答机器人,揭示检索增强与工具调用组件效能差异。
- 适合研究数据库智能问答、提示工程与系统评估的开发者与学者。
大型语言模型(LLM)的发展已革新多个行业的问答能力,包括数据库领域。然而,当前仍缺乏全面评估不同大模型及其模块化组件在数据库问答中能力的基准。为此,我们提出DQABench,首个面向大模型的综合性数据库问答基准。DQABench采用创新的基于大模型的方法,自动完成评估数据集的生成、清洗与重写,生成超过20万条中英文问答对。这些问答对涵盖从手册、在线社区及数据库实例中提取的广泛数据库知识,支持对大模型在数据库问答任务中检索增强生成(RAG)与工具调用生成(TIG)能力的额外评估。此外,我们构建了高度模块化且可扩展的DQATestbed数据库问答测试平台,包含基础与高级组件如问题分类路由(QCR)、RAG、TIG和提示模板工程(PTE)。DQABench还提供标准化评估流程,计算多种指标以确保评估的准确性与公平性。我们利用DQABench在该测试平台上全面评估大模型的数据库问答能力,发现:(i) 九个基于大模型的问答机器人各自的优劣势;(ii) 不同服务组件(如QCR、RAG、TIG)对性能的影响及改进空间。本基准与发现将为未来大模型驱动的数据库问答研究提供指导。
原文摘要 · Abstract (English)
The development of Large Language Models (LLMs) has revolutionized QA across various industries, including the database domain. However, there is still a lack of a comprehensive benchmark to evaluate the capabilities of different LLMs and their modular components in database QA. To this end, we introduce DQABench, the first comprehensive database QA benchmark for LLMs. DQABench features an innovative LLM-based method to automate the generation, cleaning, and rewriting of evaluation dataset, resulting in over 200,000 QA pairs in English and Chinese, separately. These QA pairs cover a wide range of database-related knowledge extracted from manuals, online communities, and database instances. This inclusion allows for an additional assessment of LLMs' Retrieval-Augmented Generation (RAG) and Tool Invocation Generation (TIG) capabilities in the database QA task. Furthermore, we propose a comprehensive LLM-based database QA testbed DQATestbed. This testbed is highly modular and scalable, with basic and advanced components such as Question Classification Routing (QCR), RAG, TIG, and Prompt Template Engineering (PTE). Moreover, DQABench provides a comprehensive evaluation pipeline that computes various metrics throughout a standardized evaluation process to ensure the accuracy and fairness of the evaluation. We use DQABench to evaluate the database QA capabilities under the proposed testbed comprehensively. The evaluation reveals findings like (i) the strengths and limitations of nine LLM-based QA bots and (ii) the performance impact and potential improvements of various service components (e.g., QCR, RAG, TIG). Our benchmark and findings will guide the future development of LLM-based database QA research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。