SPIO用多路径规划提升自动数据分析的准确性和可靠性
SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science
- 引入多智能体协同规划,动态生成多种数据处理策略
- 在Kaggle和OpenML上平均性能比现有方法高5.6%
- 支持单路径最优与多路径集成两种模式,适合复杂场景
大型语言模型(LLMs)已推动自动化数据分析中的动态推理,但现有多智能体系统受限于僵化的单路径工作流,难以开展策略探索,常导致次优结果。为此,我们提出SPIO(序列计划集成与优化),该框架通过四个核心模块——数据预处理、特征工程、模型选择和超参数调优——取代固定流程,采用自适应多路径规划。每个模块中,专业智能体生成多样候选策略,经优化智能体逐级串联与精炼。SPIO提供两种运行模式:SPIO-S用于选择单一最优流水线,SPIO-E则对前k个最优流水线进行集成以增强鲁棒性。在Kaggle和OpenML基准上的大量实验表明,SPIO持续优于当前最先进基线,平均性能提升5.6%。通过显式探索并融合多条解决方案路径,SPIO为自动化数据科学提供了更灵活、精准且可靠的基石。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have enabled dynamic reasoning in automated data analytics, yet recent multi-agent systems remain limited by rigid, single-path workflows that restrict strategic exploration and often lead to suboptimal outcomes. To overcome these limitations, we propose SPIO (Sequential Plan Integration and Optimization), a framework that replaces rigid workflows with adaptive, multi-path planning across four core modules: data preprocessing, feature engineering, model selection, and hyperparameter tuning. In each module, specialized agents generate diverse candidate strategies, which are cascaded and refined by an optimization agent. SPIO offers two operating modes: SPIO-S for selecting a single optimal pipeline, and SPIO-E for ensembling top-k pipelines to maximize robustness. Extensive evaluations on Kaggle and OpenML benchmarks show that SPIO consistently outperforms state-of-the-art baselines, achieving an average performance gain of 5.6%. By explicitly exploring and integrating multiple solution paths, SPIO delivers a more flexible, accurate, and reliable foundation for automated data science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。