用分层框架生成更真实表格数据,兼顾逻辑一致与统计特性。
Hierarchical Synthetic Tabular Data Generation: A Hybrid Top-Down and Bottom-Up Framework

- 分上下两路生成:上层定规则,下层学模式,再融合迭代优化。
- 在金融多模态数据上,合成数据训练效果优于纯神经网络方法。
- 适合需要高一致性与稀有事件覆盖的低数据量场景使用。
现有合成表格数据生成方法或依赖纯生成模型,或基于大语言模型,均面临数据异质性、逻辑不一致、罕见事件覆盖不足及低数据场景下鲁棒性差的问题。本文提出一种分层混合自上而下与自下而上的(H-TDBU)框架,将语义结构与随机纹理解耦。上层路径构建结构驱动的逻辑约束与跨模态对齐规则;下层路径采用轻量级表格生成器,从真实数据中学习局部统计模式。两条路径通过统一合成引擎整合,并引入迭代反馈机制。我们在结合表格与情感文本的弱多模态金融基准上进行评估,实验表明,所提H-TDBU方法在训练-合成-测试-真实性能上优于神经基线模型,同时保持了语义一致性。结果表明,分层规则引导的合成机制可有效实现可控性、语义连贯性与统计保真度的平衡。
原文摘要 · Abstract (English)
Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes. In this paper, we propose a hierarchical hybrid top-down and bottom-up (H-TDBU) framework that decouples semantic structures from stochastic texture. In the top-down path, structure-driven logical constraints and cross-modal alignment rules are constructed, while in the bottom-up path, lightweight tabular generators are used to learn local statistical patterns from real data. The two paths are consolidated in a unified synthesis engine with an iterative feedback loop. We evaluate the framework on weak multimodal financial benchmarks combining tabular and sentiment-text data. Experimental results show that our H-TDBU approach improves train-synthetic-test-real performance over neural baseline methods while preserving semantic consistency. Our results suggest that hierarchical rule-guided synthesis provides an effective mechanism for combining controllability, semantic coherence, and statistical fidelity in synthetic data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。