用合成数据精准估计模型错误率,解决小样本评估难题。
Using Synthetic Data to estimate the True Error is theoretically and practically doable
- 基于合成数据构建新的泛化界,指导优化生成样本。
- 在有限标注数据下,误差估计准确度显著优于基线方法。
- 适合资源受限场景下的模型评估,尤其适用于数据稀缺任务。
准确评估模型性能对机器学习系统在真实场景中的部署至关重要。传统方法依赖大规模标注测试集以确保评估可靠性,但在许多情况下,获取大样本标注数据成本高昂且耗时。因此,我们常需在少量标注样本条件下进行评估,这在理论上极具挑战性。近期生成模型的发展为合成高质量数据提供了可能。本文系统研究了在标注数据有限的条件下,利用合成数据估计模型测试误差的可行性。为此,我们提出了考虑合成数据的新泛化边界,这些边界揭示了生成器质量对评估的关键作用,并指导优化合成样本的生成策略。受此启发,我们提出一种理论支撑的合成数据生成方法,用于模型评估。在模拟和表格数据集上的实验表明,相比现有基线方法,该方法能更准确、更可靠地估计测试误差。
原文摘要 · Abstract (English)
Accurately evaluating model performance is crucial for deploying machine learning systems in real-world applications. Traditional methods often require a sufficiently large labeled test set to ensure a reliable evaluation. However, in many contexts, a large labeled dataset is costly and labor-intensive. Therefore, we sometimes have to do evaluation by a few labeled samples, which is theoretically challenging. Recent advances in generative models offer a promising alternative by enabling the synthesis of high-quality data. In this work, we make a systematic investigation about the use of synthetic data to estimate the test error of a trained model under limited labeled data conditions. To this end, we develop novel generalization bounds that take synthetic data into account. Those bounds suggest novel ways to optimize synthetic samples for evaluation and theoretically reveal the significant role of the generator's quality. Inspired by those bounds, we propose a theoretically grounded method to generate optimized synthetic data for model evaluation. Experimental results on simulation and tabular datasets demonstrate that, compared to existing baselines, our method achieves accurate and more reliable estimates of the test error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。