arXiv:2608.10579cs.AI2026-08

用自然语言让AI自动选数据,省去手动调规则的麻烦。

Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent

论文配图:Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent
图 1 · 摘自论文原文
  • 用户用自然语言描述需求,AI自动组合最佳数据筛选策略。
  • 在数学、医疗、编程领域表现优于传统方法,部分场景超越全量训练。
  • 适合需要快速构建高质量指令数据的开发者和研究者。

现有指令数据选择方法虽引入多种度量标准,但现实数据集的复杂性导致单一指标难以泛化。开发者常需手动检查数据并为每个新应用设计启发式规则,过程繁琐且易出错。本文提出一种范式转变:通过指令数据选择智能体DataMaster,实现从人工配置到自动化编排的升级。用户只需以自然语言描述数据需求,DataMaster即可自主理解意图并生成最优选择策略。跨数学、医疗、代码领域的大量实验表明,DataMaster在多数设置中优于静态基线,在众多场景下甚至超越全量数据训练。DataMaster的实现及复现脚本已公开于https://github.com/nju-websoft/DataMaster。

原文摘要 · Abstract (English)

Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often forced to manually inspect data and craft heuristic rules for each new application---a tedious and error-prone process. In this paper, we propose a paradigm shift from manual configuration to automated orchestration via the Instruction Data Selection Agent (DataMaster), which interprets user intent and autonomously composes optimal selection strategies. By allowing users to specify data needs through natural language descriptions, DataMaster simplifies data curation and removes the burden of manual strategy design. Extensive experiments across the math, medical, and code domains show that DataMaster outperforms static baselines in most settings and surpasses full-pool training in a substantial number of cases. The implementation of DataMaster and the scripts needed to reproduce the reported pipeline are publicly available at https://github.com/nju-websoft/DataMaster.

数据筛选智能体自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。