测试AI在垂直领域数据科学中的表现,发现人机协作胜过纯AI。
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
- 构建跨6大行业的17个任务基准,评估人机协同与纯AI表现
- 纯AI表现低于竞赛前25%人类选手,人机协作效果最佳
- 揭示当前AI在专业推理上的局限,适合关注人机协作的研究者
数据科学在多个领域将复杂数据转化为可行动洞察至关重要。尽管大型语言模型和AI代理已显著自动化数据科学流程,但其在特定领域任务上能否达到人类专家水平仍不明确,人类专长在哪些方面仍具优势亦未清晰。本文提出AgentDS,一个面向垂直领域数据科学的基准与竞赛平台,涵盖商业、食品生产、医疗、保险、制造和零售银行共17项挑战。我们组织了包含29支团队、80名参与者的开放竞赛,系统比较了人机协作与纯AI基线的表现。结果表明,当前AI代理在领域特定推理上表现不足:纯AI基线性能低于参赛者前25%水平,而最优解均来自人机协作。该发现挑战了完全自动化叙事,强调了人类专长在数据科学中的持续价值,并指明下一代AI的发展方向。访问官网:https://agentds.org/,开源数据集见:https://huggingface.co/datasets/lainmn/AgentDS。
原文摘要 · Abstract (English)
Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) and artificial intelligence (AI) agents have significantly automated data science workflow. However, it remains unclear to what extent AI agents can match the performance of human experts on domain-specific data science tasks, and in which aspects human expertise continues to provide advantages. We introduce AgentDS, a benchmark and competition designed to evaluate both AI agents and human-AI collaboration performance in domain-specific data science. AgentDS consists of 17 challenges across six industries: commerce, food production, healthcare, insurance, manufacturing, and retail banking. We conducted an open competition involving 29 teams and 80 participants, enabling systematic comparison between human-AI collaborative approaches and AI-only baselines. Our results show that current AI agents struggle with domain-specific reasoning. AI-only baselines perform below the top quartile of competition participants, while the strongest solutions arise from human-AI collaboration. These findings challenge the narrative of complete automation by AI and underscore the enduring importance of human expertise in data science, while illuminating directions for the next generation of AI. Visit the AgentDS website here: https://agentds.org/ and open source datasets here: https://huggingface.co/datasets/lainmn/AgentDS .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。