arXiv:2603.10823stat.MLcs.LG2026-03

用强化学习让生成数据更贴近真实预测信号,提升小样本场景下模型表现

ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement Learning

  • 通过强化学习直接优化特征相关性保留,引导生成器聚焦关键预测特征
  • 在小样本、类别不平衡等挑战下,下游任务准确率超越现有方法
  • 支持专家约束控制,适合隐私保护与数据稀缺场景的数据合成

深度生成模型可通过生成合成数据缓解数据稀缺与隐私问题,但在低数据、类别不平衡的表格数据场景中,难以充分学习复杂的数据分布。我们认为,追求完整的联合分布可能过度复杂;基于最新理论分析,应优先学习条件分布 $P(yackslashmid \bm{X})$ 以提高数据效率。为此,我们提出 extbf{ReTabSyn}——一种基于强化学习的表格数据合成框架,在生成器训练过程中提供关于特征相关性保持的直接反馈,促使生成器在数据有限时优先学习最具预测价值的信号,从而增强下游模型性能。我们基于语言模型构建生成器,并在小样本、类别不平衡及分布偏移等多种基准上进行实证验证,结果表明 ReTabSyn 持续优于当前最优基线。此外,该方法可轻松扩展以控制合成数据的多种属性,如施加专家指定的约束条件于生成样本。

原文摘要 · Abstract (English)

Deep generative models can help with data scarcity and privacy by producing synthetic training data, but they struggle in low-data, imbalanced tabular settings to fully learn the complex data distribution. We argue that striving for the full joint distribution could be overkill; for greater data efficiency, models should prioritize learning the conditional distribution $P(y\mid \bm{X})$, as suggested by recent theoretical analysis. Therefore, we overcome this limitation with \textbf{ReTabSyn}, a \textbf{Re}inforced \textbf{Tab}ular \textbf{Syn}thesis pipeline that provides direct feedback on feature correlation preservation during synthesizer training. This objective encourages the generator to prioritize the most useful predictive signals when training data is limited, thereby strengthening downstream model utility. We empirically fine-tune a language model-based generator using this approach, and across benchmarks with small sample sizes, class imbalance, and distribution shift, ReTabSyn consistently outperforms state-of-the-art baselines. Moreover, our approach can be readily extended to control various aspects of synthetic tabular data, such as applying expert-specified constraints on generated observations.

表格生成强化学习数据合成小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。