arXiv:2509.09960cs.LGcs.AI2025-09

用规则引导生成,解决小数据下表格合成的偏差与冗余问题。

Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

  • 提取可解释模型的条件规则,嵌入提示词引导生成符合领域分布的数据。
  • 双粒度过滤机制有效减少重复采样,保留稀有但关键样本。
  • 在极低数据场景下平均性能提升7.48%,适合小样本领域应用。

合成表格数据在机器学习中日益重要,尤其在真实高质量数据不足时支持下游任务。现有方法如生成对抗网络(GANs)和微调的大语言模型(LLMs)通常需要充足参考数据,限制了其在记录稀少的特定领域数据集中的应用。尽管基于提示的LLMs无需参数调优,但常生成分布漂移且局部重复的数据,导致下游任务性能下降。为此,我们提出ReFine框架:(i) 从可解释模型中提取if-then规则,并将其嵌入提示词,显式引导生成过程贴近领域分布;(ii) 采用双粒度过滤机制,缓解过度采样模式,同时保留稀有但信息丰富的样本,降低局部冗余。在多种基准上的实验表明,ReFine具备稳健的下游效用,在不同数据集和数据环境下平均排名顶尖,极端低数据场景下平均相对提升达7.48%。

原文摘要 · Abstract (English)

Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient. Existing tabular generation approaches, such as generative adversarial networks (GANs) and fine-tuned Large Language Models (LLMs), typically require sufficient reference data, limiting their effectiveness in domain-specific datasets with scarce records. While prompt-based LLMs offer flexibility without parameter tuning, they often generate distributionally drifted data with localized redundancy, leading to degradation in downstream task performance. To overcome these issues, we propose ReFine, a framework that (i) extracts symbolic if-then rules from interpretable models and embeds them into prompts to explicitly guide the generation process toward the domain-specific distribution, and (ii) applies dual-granularity filtering that mitigates over-sampling patterns while preserving rare but informative samples to reduce localized redundancy. Extensive experiments on diverse benchmarks demonstrate that ReFine provides robust downstream utility, achieving a top-tier average rank across datasets and data regimes, with an average relative improvement of 7.48% in extreme low-data regimes.

表格生成小样本大模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。