根据数据压力特征自动选最合适的表格生成模型,提升实用性和效果。
SYNTHONY: A Stress-Aware, Intent-Conditioned Agent for Deep Tabular Generative Models Selection
- 通过四个可解释的压力维度量化数据难度,构建元特征
- 在7个数据集上实现超过随机选择的顶尖选型准确率
- 适合需要平衡真实性、隐私与可用性的数据合成场景
面向表格数据的深度生成模型(GAN、扩散模型、基于LLM的生成器)在不同数据集上表现差异巨大;最佳生成器家族高度依赖于长尾分布、高基数分类变量、齐普夫不均衡和小样本等分布性压力。这种脆弱性使得实际部署困难,尤其当用户需权衡保真度、隐私和效用时。本文研究意图条件下的表格生成器选择:给定数据集和用户对评估指标的偏好,目标是选出相对于特定意图最优的生成器。提出「压力画像」方法,构建针对生成器的元特征表示,以量化数据在四个可解释压力维度上的难度,并集成至「SYNTHONY」选型框架中,该框架将压力画像与校准过的生成器能力注册表匹配。在涵盖7个数据集、10种生成器、3种意图的基准测试中,基于压力的元特征对生成器性能具有极强预测力:使用kNN的选择器达到优异的Top-1选型准确率,显著优于零样本大模型选择器和随机基线。分析发现,元特征选择与能力选择之间的差距主要源于人工设计的能力注册表,提示未来应探索学习式能力表示。
原文摘要 · Abstract (English)
Deep generative models for tabular data (GANs, diffusion models, and LLM-based generators) exhibit highly non-uniform behavior across datasets; the best-performing synthesizer family depends strongly on distributional stressors such as long-tailed marginals, high-cardinality categorical, Zipfian imbalance, and small-sample regimes. This brittleness makes practical deployment challenging, especially when users must balance competing objectives of fidelity, privacy, and utility. We study {intent-conditioned tabular synthesis selection}: given a dataset and a user intent expressed as a preference over evaluation metrics, the goal is to select a synthesizer that minimizes regret relative to an intent-specific oracle. We propose {stress profiling}, a synthesis-specific meta-feature representation that quantifies dataset difficulty along four interpretable stress dimensions, and integrate it into {SYNTHONY}, a selection framework that matches stress profiles against a calibrated capability registry of synthesizer families. Across a benchmark of 7 datasets, 10 synthesizers, and 3 intents, we demonstrate that stress-based meta-features are highly predictive of synthesizer performance: a $k$NN selector using these features achieves strong Top-1 selection accuracy, substantially outperforming zero-shot LLM selectors and random baselines. We analyze the gap between meta-feature-based and capability-based selection, identifying the hand-crafted capability registry as the primary bottleneck and motivating learned capability representations as a direction for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。