Flowco让非程序员也能高效编写可调试的数据分析流程
Flowco: Rethinking Data Analysis in the Age of LLMs
- 用可视化数据流+LLM,全程辅助代码生成与修改
- 用户研究显示,新手能更快完成分析并修复错误
- 适合需要精细控制和反复迭代的科研/业务分析者
传统数据分析依赖编程实现数据转换、可视化、分析与解释。大语言模型(LLMs)现可生成简单分析的代码,有望降低数据科学门槛,使无编程背景者也能开展科研、商业与政策分析。然而,在真实场景中,分析师常需对特定步骤精细控制,显式验证中间结果,并迭代优化分析方法。这些需求使得仅靠LLM或现有工具(如计算笔记本)难以构建稳健且可复现的分析流程。本文提出Flowco,一种新型人机协同系统,采用可视化数据流编程范式,并将LLM深度融入整个创作过程。用户研究表明,Flowco显著提升了分析师(尤其是编程经验较少者)在快速编写、调试和优化数据分析流程方面的能力。
原文摘要 · Abstract (English)
Conducting data analysis typically involves authoring code to transform, visualize, analyze, and interpret data. Large language models (LLMs) are now capable of generating such code for simple, routine analyses. LLMs promise to democratize data science by enabling those with limited programming expertise to conduct data analyses, including in scientific research, business, and policymaking. However, analysts in many real-world settings must often exercise fine-grained control over specific analysis steps, verify intermediate results explicitly, and iteratively refine their analytical approaches. Such tasks present barriers to building robust and reproducible analyses using LLMs alone or even in conjunction with existing authoring tools (e.g., computational notebooks). This paper introduces Flowco, a new mixed-initiative system to address these challenges. Flowco leverages a visual dataflow programming model and integrates LLMs into every phase of the authoring process. A user study suggests that Flowco supports analysts, particularly those with less programming experience, in quickly authoring, debugging, and refining data analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。