GRACE-DS评估大模型自动机器学习代理在真实数据科学流程中的表现。
GRACE-DS: a Guarded Reward-guided Agent Correction Environment in Data Science

- 构建隔离环境,模拟从数据探索到代码修复的全流程
- 相比单次生成等基线,端到端测试质量提升且更符合规范
- 适用于组织定制化需求,验证模型合规性与可复现性
我们提出GRACE-DS,一个用于预部署评估大语言模型驱动的AutoML代理的数据科学受控评估环境。该环境提供一套评估指标,适用于特定组织的表格型机器学习任务。它让代理经历从规划、数据检查、特征工程、模型开发、验证到代码修复和最终提交的完整工作流,同时隐藏的可执行验证器不仅评估最终预测性能,还检测数据泄露、可复现性、协议有效性、修正行为与奖励对齐。在超过7000个实验中验证,采用灵活迭代交互(本文方法)的结构化策略,在端到端归一化隐测质量上优于单次生成、非结构化交互及重启基线,并显著提升协议有效完成率。结果表明,GRACE-DS是评估大模型型AutoML代理在类生产环境中执行机器学习工作流能力的可靠平台,且符合组织特定要求。
原文摘要 · Abstract (English)
We introduce GRACE-DS, a Guarded Reward-guided Agent Correction Environment in Data Science for pre-deployment evaluation of LLM-powered AutoML agents. GRACE-DS is a set of evaluation metrics in an isolated environment that can be applied to tabular ML tasks specific to a particular organization. It exposes agents to realistic workflow stages, from planning and data inspection through feature engineering, model development, validation, and code repair to final submission, while hidden executable validators measure not only final predictive performance but also leakage avoidance, reproducibility, protocol validity, correction behavior, and reward alignment. The strongest structured regime, flexible iterative interaction (our approach), achieves higher end-to-end normalized hidden-test quality than single-shot generation, unstructured interaction, and restart-based baselines, while also improving protocol-valid completion. Validated across more than 7,000 episodes, these results establish GRACE-DS as a robust platform for assessing the capacity of LLM-based AutoML agents to execute machine learning workflows under production-like conditions and in accordance with organization-specific requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。