arXiv:2606.25996cs.AIcs.CL2026-06被引 7

AI代理自动生成高质量合成数据,让数据生产像模型训练一样可迭代优化。

Autodata: An agentic data scientist to create high quality synthetic data

论文配图:Autodata: An agentic data scientist to create high quality synthetic data
图 1 · 摘自论文原文
  • 用元优化训练AI代理扮演数据科学家,自动构建优质数据集。
  • 在代码、法律推理和数学推理任务上,合成数据效果优于传统方法。
  • 适合希望提升数据质量、探索自动化数据生成的研究者和工程师。

我们提出Autodata,一种通用方法,使AI代理能像数据科学家一样构建高质量的训练与评估数据。通过元优化训练此类数据科学家代理,使其学会生成更优的数据。文中描述了整体框架及具体实现方式——Agentic Self-Instruct。我们在计算机科学研究、法律推理与数学对象推理任务上进行实验,结果表明其合成数据性能优于经典方法。进一步对数据科学家代理本身进行元优化,带来更大的性能提升。该方法为将更多推理算力转化为更高品质的模型训练数据提供了新路径。总体而言,这一方向有望改变AI数据的构建方式。

原文摘要 · Abstract (English)

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning with mathematical objects, where we obtain improved results compared to classical synthetic dataset creation methods. Further, meta-optimizing the data scientist agent itself delivers an even larger performance uplift. Agentic data creation provides a way to convert increased inference compute into higher quality model training. Overall, we believe this direction has the potential to change the way we build AI data.

合成数据AI代理元学习数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。