构建首个流程挖掘领域中英双语Text-to-SQL数据集
Text-to-SQL Oriented to the Process Mining Domain: A PT-EN Dataset for Query Translation
- 聚焦流程挖掘,设计含专业术语的中英双语数据集
- 包含1655条自然语言查询与205条对应SQL语句
- 适合流程挖掘与自然语言转SQL研究者使用
本文提出text-2-SQL-4-PM,一个面向流程挖掘领域的中英双语基准数据集,用于文本转SQL任务。该数据集针对流程挖掘特有的专业词汇和事件日志衍生的单表结构设计,包含1655条自然语言语句(含人工改写)、205条SQL语句及10个限定条件。通过专家手工筛选、专业翻译与详尽标注流程确保质量。基于GPT-3.5 Turbo的基线实验验证了其可用性,表明该数据集可支持文本转SQL模型评估,并拓展至语义解析等自然语言处理任务。
原文摘要 · Abstract (English)
This paper introduces text-2-SQL-4-PM, a bilingual (Portuguese-English) benchmark dataset designed for the text-to-SQL task in the process mining domain. Text-to-SQL conversion facilitates natural language querying of databases, increasing accessibility for users without SQL expertise and productivity for those that are experts. The text-2-SQL-4-PM dataset is customized to address the unique challenges of process mining, including specialized vocabularies and single-table relational structures derived from event logs. The dataset comprises 1,655 natural language utterances, including human-generated paraphrases, 205 SQL statements, and ten qualifiers. Methods include manual curation by experts, professional translations, and a detailed annotation process to enable nuanced analyses of task complexity. Additionally, a baseline study using GPT-3.5 Turbo demonstrates the feasibility and utility of the dataset for text-to-SQL applications. The results show that text-2-SQL-4-PM supports evaluation of text-to-SQL implementations, offering broader applicability for semantic parsing and other natural language processing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。