用科学原则指导的智能数据科学系统,提升自动化流程可信度。
VDSAgents: A PCS-Guided Multi-Agent System for Veridical Data Science Automation
- 基于预测性-可计算性-稳定性原则,分阶段协同处理数据任务。
- 在9个数据集上优于AutoKaggle和DataInterpreter,性能更稳定。
- 适合追求可解释、可审计数据科学结果的研究者与工程师。
大型语言模型(LLMs)正越来越多地融入数据科学自动化流程。然而,这些由LLM驱动的系统仅依赖其内部推理,缺乏来自科学与理论原则的指导,限制了其在噪声大、结构复杂的现实数据上的可信度与鲁棒性。本文提出VDSAgents,一个基于验证性数据科学(VDS)框架中提出的预测性-可计算性-稳定性(PCS)原则的多智能体系统。该系统采用模块化工作流,涵盖数据清洗、特征工程、建模与评估,每个阶段均由专门智能体处理,并结合扰动分析、单元测试与模型验证,确保功能正确性与科学可审计性。我们在九个具有不同特性的数据集上进行评估,使用DeepSeek-V3和GPT-4o作为后端,对比当前最先进的端到端数据科学系统如AutoKaggle和DataInterpreter。结果表明,VDSAgents在各项指标上持续优于基线方法,验证了将PCS原则嵌入LLM驱动的数据科学自动化的可行性。
原文摘要 · Abstract (English)
Large language models (LLMs) become increasingly integrated into data science workflows for automated system design. However, these LLM-driven data science systems rely solely on the internal reasoning of LLMs, lacking guidance from scientific and theoretical principles. This limits their trustworthiness and robustness, especially when dealing with noisy and complex real-world datasets. This paper provides VDSAgents, a multi-agent system grounded in the Predictability-Computability-Stability (PCS) principles proposed in the Veridical Data Science (VDS) framework. Guided by PCS principles, the system implements a modular workflow for data cleaning, feature engineering, modeling, and evaluation. Each phase is handled by an elegant agent, incorporating perturbation analysis, unit testing, and model validation to ensure both functionality and scientific auditability. We evaluate VDSAgents on nine datasets with diverse characteristics, comparing it with state-of-the-art end-to-end data science systems, such as AutoKaggle and DataInterpreter, using DeepSeek-V3 and GPT-4o as backends. VDSAgents consistently outperforms the results of AutoKaggle and DataInterpreter, which validates the feasibility of embedding PCS principles into LLM-driven data science automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。