AutoKaggle用多智能体协作自动完成数据科学任务,提升效率与准确率。
AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
- 多智能体协同执行数据清洗、特征工程与建模流程
- 在8个真实竞赛中验证,提交成功率0.85,综合得分0.82
- 支持人工干预,兼顾自动化与专家经验
涉及表格数据的数据科学任务面临复杂挑战,需要高级求解策略。我们提出AutoKaggle,一个强大且以用户为中心的框架,通过协作式多智能体系统辅助数据科学家完成日常数据管道工作。该框架采用迭代开发流程,结合代码执行、调试与全面单元测试,确保代码正确性与逻辑一致性。提供高度可定制的工作流,允许用户在每个阶段介入,实现自动化智能与人类专长的融合。其通用数据科学工具包包含经验证的数据清洗、特征工程和建模函数,显著提升常见任务的处理效率。我们在8个Kaggle竞赛中模拟真实应用场景下的数据处理流程。评估结果表明,AutoKaggle在典型数据科学流水线中达到0.85的验证提交率和0.82的综合得分,充分证明其在处理复杂数据科学任务中的有效性与实用性。
原文摘要 · Abstract (English)
Data science tasks involving tabular data present complex challenges that require sophisticated problem-solving approaches. We propose AutoKaggle, a powerful and user-centric framework that assists data scientists in completing daily data pipelines through a collaborative multi-agent system. AutoKaggle implements an iterative development process that combines code execution, debugging, and comprehensive unit testing to ensure code correctness and logic consistency. The framework offers highly customizable workflows, allowing users to intervene at each phase, thus integrating automated intelligence with human expertise. Our universal data science toolkit, comprising validated functions for data cleaning, feature engineering, and modeling, forms the foundation of this solution, enhancing productivity by streamlining common tasks. We selected 8 Kaggle competitions to simulate data processing workflows in real-world application scenarios. Evaluation results demonstrate that AutoKaggle achieves a validation submission rate of 0.85 and a comprehensive score of 0.82 in typical data science pipelines, fully proving its effectiveness and practicality in handling complex data science tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。