AI在数据管道自动化中能力被低估,因基准测试存在错误
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
- 用改进的LLM和人工验证方法审计基准测试质量
- 修正后模型成功率显著提升,证明原结果受错误评估误导
- 适合关注AI自动化、数据工程评估可靠性的研究者
构建提取-加载-转换(ELT)数据管道是劳动密集型的数据工程任务,也是AI自动化的高价值目标。在首个端到端ELT管道构建基准测试ELT-Bench上,早期AI代理表现不佳,暗示其缺乏实用性。本文重新审视这些结果,发现两个导致能力严重低估的因素:首先,使用升级版大语言模型重新评估表明,提取与加载阶段已基本解决,转换性能显著提升;其次,我们提出审计-修正方法,结合可扩展的LLM根因分析与严格的人工验证(组间一致性Fleiss' kappa = 0.85),对基准测试质量进行审计。应用该方法于ELT-Bench,发现多数失败的转换任务源于基准本身的问题——包括僵化评价脚本、模糊规格说明和错误真值——这些错误惩罚了正确的代理输出。基于此,我们构建了修复后的ELT-Bench-Verified,包含优化的评估逻辑和修正的真值。在该版本上重新评估显示,性能提升完全归因于基准修正。结果表明,模型快速进步与基准质量问题共同导致了对代理能力的低估。更广泛地,这一发现呼应了文本到SQL基准中普遍存在的标注错误,提示数据工程评估中的质量问题是系统性问题。应将系统性质量审计作为复杂智能体任务的标准实践。我们公开发布ELT-Bench-Verified,以提供更可靠的进展基础。
原文摘要 · Abstract (English)
Constructing Extract-Load-Transform (ELT) pipelines is a labor-intensive data engineering task and a high-impact target for AI automation. On ELT-Bench, the first benchmark for end-to-end ELT pipeline construction, AI agents initially showed low success rates, suggesting they lacked practical utility. We revisit these results and identify two factors causing a substantial underestimation of agent capabilities. First, re-evaluating ELT-Bench with upgraded large language models reveals that the extraction and loading stage is largely solved, while transformation performance improves significantly. Second, we develop an Auditor-Corrector methodology that combines scalable LLM-driven root-cause analysis with rigorous human validation (inter-annotator agreement Fleiss' kappa = 0.85) to audit benchmark quality. Applying this to ELT-Bench uncovers that most failed transformation tasks contain benchmark-attributable errors -- including rigid evaluation scripts, ambiguous specifications, and incorrect ground truth -- that penalize correct agent outputs. Based on these findings, we construct ELT-Bench-Verified, a revised benchmark with refined evaluation logic and corrected ground truth. Re-evaluating on this version yields significant improvement attributable entirely to benchmark correction. Our results show that both rapid model improvement and benchmark quality issues contributed to underestimating agent capabilities. More broadly, our findings echo observations of pervasive annotation errors in text-to-SQL benchmarks, suggesting quality issues are systemic in data engineering evaluation. Systematic quality auditing should be standard practice for complex agentic tasks. We release ELT-Bench-Verified to provide a more reliable foundation for progress in AI-driven data engineering automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。