LLM改变数据标注方式,从经验驱动转向洞察导向。
The Evolution of LLM Adoption in Industry Data Curation Practices
- 用大模型生成辅助数据集,补充传统人工标注。
- 2024年12名从业者测试显示,新流程提升数据理解效率。
- 适合想优化数据标注的工程师和产品团队看。
随着大语言模型(LLMs)在处理非结构化文本数据方面能力不断增强,它们为改进数据整理工作流提供了新机遇。本文通过一系列调查、访谈和用户研究,探讨了一家大型科技公司实践中LLM应用的演变。2023年第二季度,我们对84名开发人员进行了调查,评估其在开发任务中的LLM使用情况;第三季度对10位专家进行访谈,了解数据需求变化。2024年第二季度,通过包含两个基于LLM原型的用户研究(参与者N=12),探索从业者当前及未来对LLM的使用预期。研究发现,数据理解模式正从以经验为主导的自下而上方法,转向由洞察驱动的自上而下流程。为应对更复杂的数据环境,数据从业者现在以大模型生成的‘银级’数据集补充传统由领域专家创建的‘金级’数据集,并通过多方专家验证构建更高标准的‘超级金级’数据集。该研究揭示了LLMs在大规模非结构化数据分析中的变革性作用,也为工具进一步发展指明方向。
原文摘要 · Abstract (English)
As large language models (LLMs) grow increasingly adept at processing unstructured text data, they offer new opportunities to enhance data curation workflows. This paper explores the evolution of LLM adoption among practitioners at a large technology company, evaluating the impact of LLMs in data curation tasks through participants' perceptions, integration strategies, and reported usage scenarios. Through a series of surveys, interviews, and user studies, we provide a timely snapshot of how organizations are navigating a pivotal moment in LLM evolution. In Q2 2023, we conducted a survey to assess LLM adoption in industry for development tasks (N=84), and facilitated expert interviews to assess evolving data needs (N=10) in Q3 2023. In Q2 2024, we explored practitioners' current and anticipated LLM usage through a user study involving two LLM-based prototypes (N=12). While each study addressed distinct research goals, they revealed a broader narrative about evolving LLM usage in aggregate. We discovered an emerging shift in data understanding from heuristic-first, bottom-up approaches to insights-first, top-down workflows supported by LLMs. Furthermore, to respond to a more complex data landscape, data practitioners now supplement traditional subject-expert-created 'golden datasets' with LLM-generated 'silver' datasets and rigorously validated 'super golden' datasets curated by diverse experts. This research sheds light on the transformative role of LLMs in large-scale analysis of unstructured data and highlights opportunities for further tool development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。