arXiv:2509.21479cs.LG2025-09被引 3

用置信度过滤合成数据,提升生成质量并控制风险。

Filtering with Confidence: When Data Augmentation Meets Conformal Prediction

  • 结合置信区间预测,筛选高质量合成数据。
  • 在多个任务中提升F1分数最高达40个百分点。
  • 无需模型内部参数,适合资源有限的场景。

在众多应用中表现出色,合成数据增强成为缓解数据稀缺和应对日益依赖数据模型的有效方案。其有效性在于以降低估计器方差的方式扩展训练集,同时引入极小偏差。因此控制偏差至关重要:有效的数据增强应从与训练集相同的基础分布生成多样化样本,且偏差最小。本文提出置信度数据增强,一种基于置信区间预测的原理性数据过滤框架,可在生成多样化合成数据的同时,以可证明的风险控制方式剔除低质量生成。该方法实现简单,无需访问模型内部对数几率,也无需大规模重训练。我们在主题预测、情感分析、图像分类和欺诈检测等多个任务上验证了该方法的有效性,结果表明,相比未增强基线,F1得分最高提升40个百分点;相较于其他过滤增强基线,仍能提升4~6个百分点。

原文摘要 · Abstract (English)

With promising empirical performance across a wide range of applications, synthetic data augmentation appears a viable solution to data scarcity and the demands of increasingly data-intensive models. Its effectiveness lies in expanding the training set in a way that reduces estimator variance while introducing only minimal bias. Controlling this bias is therefore critical: effective data augmentation should generate diverse samples from the same underlying distribution as the training set, with minimal shifts. In this paper, we propose conformal data augmentation, a principled data filtering framework that leverages the power of conformal prediction to produce diverse synthetic data while filtering out poor-quality generations with provable risk control. Our method is simple to implement, requires no access to internal model logits, nor large-scale model retraining. We demonstrate the effectiveness of our approach across multiple tasks, including topic prediction, sentiment analysis, image classification, and fraud detection, showing consistent performance improvements of up to 40 percentage points (pp) in $F_1$ score over unaugmented baselines, and 4~pp over other filtered augmentation baselines.

数据增强置信度合成数据风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。