让生成数据更实用:根据模型需求动态调整生成内容和时机
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting

- 用策略引导扩散模型生成更有利于提升性能的数据
- 在数据稀缺下分类准确率最高提升15.6%,回归误差降低32%
- 适合需要高效增广的机器学习初学者与工业落地场景
表格数据生成在数据稀缺领域极具吸引力,但当前方法过度关注分布保真度,难以提升下游模型性能。本文揭示了保真度与实用性之间的差距:生成目标强调分布合理性,而增广成功的关键在于所注入样本能否降低当前学习器的验证损失。为此,我们提出 TAP(Tabular Augmentation Policy),将扩散插值与轻量级、依赖学习器状态的策略结合,引导生成聚焦高价值区域,并通过显式门控和保守窗口机制控制安全注入。在七组真实世界数据集上,面对严重数据稀缺时,TAP 持续优于多个强基线,分类准确率最高提升15.6个百分点,回归均方根误差最多降低32%。
原文摘要 · Abstract (English)
Generative tabular augmentation is appealing in data-scarce domains, yet the prevailing focus on distributional fidelity does not reliably translate into better downstream models. We formalize a fidelity-utility gap: common generative objectives prioritize distributional plausibility, whereas augmentation succeeds only when injected samples reduce the current learner's held-out evaluation loss. This gap motivates learning not just how to generate, but what to generate and when to inject as training evolves. We propose TAP (Tabular Augmentation Policy), which couples diffusion inpainting with a lightweight, learner-conditioned policy to steer generation toward high-utility regions and controls safe injection via explicit gating and conservative windowed commitment. Under severe data scarcity, TAP consistently outperforms strong generative baselines on seven real-world datasets, improving classification accuracy by up to 15.6 percentage points and reducing regression RMSE by up to 32%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。