测试大模型处理数据准备任务的能力,发现其表现仍有不足。
Lost in the Pipeline: How Well Do Large Language Models Handle Data Preparation?
- 用差质量数据测试大模型的数据分析与清洗能力
- 大模型在数据准备上不如传统工具可靠
- 适合关注大模型实际应用局限的研究者
大型语言模型最近展现出在支持和自动化各类任务方面的卓越能力。本文聚焦于数据准备——数据驱动流程中关键但繁琐的步骤,探究大模型能否有效辅助用户完成数据选择与自动化处理。研究采用通用型和微调过的表格型大模型,以低质量数据集为输入,评估其在数据剖析与清洗等任务中的表现,并与传统数据准备工具进行对比。为衡量大模型能力,研究设计并验证了一个定制化质量评估模型,通过用户研究获取从业者期望,从而深入理解大模型在实际场景中的适用性与局限。
原文摘要 · Abstract (English)
Large language models have recently demonstrated their exceptional capabilities in supporting and automating various tasks. Among the tasks worth exploring for testing large language model capabilities, we considered data preparation, a critical yet often labor-intensive step in data-driven processes. This paper investigates whether large language models can effectively support users in selecting and automating data preparation tasks. To this aim, we considered both general-purpose and fine-tuned tabular large language models. We prompted these models with poor-quality datasets and measured their ability to perform tasks such as data profiling and cleaning. We also compare the support provided by large language models with that offered by traditional data preparation tools. To evaluate the capabilities of large language models, we developed a custom-designed quality model that has been validated through a user study to gain insights into practitioners' expectations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。