arXiv:2605.11231cs.LGcs.AI2026-05

精准生成填补数据空白的合成样本,提升模型性能。

LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection

论文配图:LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection
图 1 · 摘自论文原文
  • 结合边界距离与不确定性,筛选高价值合成数据。
  • 在稀疏区域生成有效样本,准确率超越传统方法。
  • 适合数据稀缺场景下的小样本学习与增量训练。

合成数据的有效性取决于其能否填补下游任务关键的训练分布空白。本文提出LiBaGS,一种轻量级、无需依赖生成器的目标化合成数据选择方法。通过融合决策边界邻近度、预测不确定性、真实数据密度及支持有效性对候选样本评分,确保所选样本既具信息量又位于真实数据流形上。采用边界间隙分配策略,聚焦稀疏但真实的决策边界邻域,而非简单增加数据或仅选最不确定样本。此外,引入边际价值停止规则判断何时停止添加合成样本,在模糊边界处赋予软标签,并使用多样性目标避免重复选择。实验表明,LiBaGS在准确率上优于经典过采样、硬增强、不确定性与密度消融实验,以及目标生成选择标准。

原文摘要 · Abstract (English)

Synthetic data is useful only when the added samples fill missing parts of the training distribution that matter for the downstream task. We introduce LiBaGS, a lightweight, generator-agnostic method for targeted synthetic training data selection. LiBaGS scores candidate synthetic samples by combining decision-boundary proximity, predictive uncertainty, real-data density, and support validity, so that selected samples are both informative and likely to remain on the real data manifold. We then use a boundary-gap allocation rule that targets sparse but realistic decision-boundary neighborhoods, rather than simply adding more data or selecting only the most uncertain candidates. LiBaGS also learns when enough synthetic samples have been added through a marginal-value stopping rule, assigns softer labels near ambiguous boundaries, and uses a diversity objective to avoid redundant near-duplicate selections. Experiments show that LiBaGS improves accuracy over classical oversampling, hard augmentation, uncertainty and density ablations, and targeted-generation selection criteria.

合成数据数据增强小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。