AI在数据科学中仍难超越人类,因缺乏领域知识。
Can Agentic AI Match the Performance of Human Data Scientists?
- 用隐藏图像变量的保险数据集测试AI与人类表现
- AI仅靠通用代码无法识别关键隐变量,准确率低
- 适合关注AI局限性与领域知识融合的研究者
数据科学在多个领域将复杂数据转化为可行动洞察至关重要。尽管大语言模型(LLMs)显著自动化了数据科学流程,但核心问题仍未解决:依赖领域知识的人类数据科学家能否被代理型AI系统超越?我们通过设计一个预测任务来探索该问题,其中关键潜在变量隐藏于相关图像数据中,而非表格特征。因此,仅生成通用表格分析代码的代理型AI表现不佳,而人类专家能利用领域知识识别出重要隐变量。我们在一个合成的房产保险数据集上验证了这一现象。实验表明,依赖通用分析流程的代理型AI不及结合领域知识的方法。这揭示了当前代理型AI在数据科学中的关键局限,并强调未来研究需开发能更好识别和整合领域知识的代理型AI系统。
原文摘要 · Abstract (English)
Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significantly automated data science workflows, but a fundamental question persists: Can these agentic AI systems truly match the performance of human data scientists who routinely leverage domain-specific knowledge? We explore this question by designing a prediction task where a crucial latent variable is hidden in relevant image data instead of tabular features. As a result, agentic AI that generates generic codes for modeling tabular data cannot perform well, while human experts could identify the important hidden variable using domain knowledge. We demonstrate this idea with a synthetic dataset for property insurance. Our experiments show that agentic AI that relies on generic analytics workflow falls short of methods that use domain-specific insights. This highlights a key limitation of the current agentic AI for data science and underscores the need for future research to develop agentic AI systems that can better recognize and incorporate domain knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。