基于学习者最近发展区理论,动态筛选模型最适训练样本
ZPD Detector: Data Selection via Capability-Difficulty Alignment for Large Language Models
- 用项目反应理论估算模型能力,匹配样本难度
- 动态识别各训练阶段最有效样本,提升数据利用效率
- 适合数据稀缺下大模型高效训练的研究者
随着大语言模型训练成本持续上升,高质量训练数据日益稀缺,如何在有限数据预算下选择高价值样本或合成有效数据已成为关键研究问题。现有数据选择方法多依赖静态标准(如难度、不确定性或启发式规则),未能建模模型与数据间动态关系。受最近发展区(ZPD)教育理论启发,我们提出ZPD Detector,一种双向建模模型与数据关系的数据选择框架,通过显式建模样本难度与模型当前能力的对齐关系,融合难度校准、基于项目反应理论(IRT)的模型能力估计及能力-难度匹配得分,动态识别每个学习阶段最具信息量的样本,显著提升数据利用效率;该动态匹配策略也为训练策略设计提供了新视角。所有代码与数据将在论文录用后公开,以支持可复现研究。
原文摘要 · Abstract (English)
As the cost of training large language models continues to increase and high-quality training data become increasingly scarce, selecting high-value samples or synthesizing effective training data under limited data budgets has emerged as a critical research problem. Most existing data selection methods rely on static criteria, such as difficulty, uncertainty, or heuristics, and fail to model the evolving relationship between the model and the data. Inspired by the educational theory of the Zone of Proximal Development (ZPD), we propose ZPD Detector, a data selection framework that adopts a bidirectional perspective between models and data by explicitly modeling the alignment between sample difficulty and the model's current capability. ZPD Detector integrates difficulty calibration, model capability estimation based on Item Response Theory (IRT), and a capability-difficulty matching score to dynamically identify the most informative samples at each learning stage, improving data utilization efficiency; moreover, this dynamic matching strategy provides new insights into training strategy design. All code and data will be released after our work be accepted to support reproducible researc
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。