Dataforge用AI自动完成数据清洗与特征优化,提升科研数据处理效率。
Dataforge: Agentic Platform for Autonomous Data Engineering
- 基于大模型的智能代理,自动执行数据清洗和特征变换
- 在多个表格数据集上表现最优,准确率显著提升
- 无需专家干预,适合科研人员快速构建高质量数据
材料发现、分子建模和气候科学等领域的AI应用日益增长,但数据准备成为关键瓶颈。来自多源的原始数据需经过清洗、归一化和转换才能用于AI训练,其中有效的特征变换与选择对模型性能至关重要。我们提出Dataforge,一个面向表格数据的大型语言模型驱动的智能数据工程平台,具备自动化、安全性和非专业用户友好性。该平台能自主执行数据清洗,并在预算约束的反馈循环中迭代优化特征操作,实现自动终止。在多个表格数据基准测试中,其下游性能达到最佳;消融实验进一步验证了路由机制、迭代优化及上下文锚定对精度与可靠性的贡献。Dataforge为构建自主数据智能体提供了可行路径,实现从原始数据到更优数据的转化。
原文摘要 · Abstract (English)
The growing demand for artificial intelligence (AI) applications in materials discovery, molecular modeling, and climate science has made data preparation a critical but labor-intensive bottleneck. Raw data from diverse sources must be cleaned, normalized, and transformed to become AI-ready, where effective feature transformation and selection are essential for robust learning. We present Dataforge, an LLM-powered agentic data engineering platform for tabular data that is automatic, safe, and non-expert friendly. It autonomously performs data cleaning and iteratively optimizes feature operations under a budgeted feedback loop with automatic stopping. Across tabular benchmarks, it achieves the best overall downstream performance; ablations further confirm the roles of routing/iterative refinement and grounding in accuracy and reliability. Dataforge demonstrates a practical path toward autonomous data agents that transform raw data from data to better data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。